Papers with natural language understanding

300 papers
CogKTR: A Knowledge-Enhanced Text Representation Toolkit for Natural Language Understanding (2022.emnlp-demos)

Copied to clipboard

Challenge: Existing knowledge-enhanced methods are limited to knowledge-intensive tasks.
Approach: They propose a knowledge-enhanced text representation toolkit for natural language understanding . it combines knowledge acquisition, knowledge representation, knowledge injection and knowledge application .
Outcome: The proposed toolkit supports knowledge acquisition, knowledge representation, knowledge injection, and knowledge application.
K-PLUG: Knowledge-injected Pre-trained Language Model for Natural Language Understanding and Generation in E-Commerce (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing pre-trained language models are not explicitly aware of domain-specific knowledge, which is essential for downstream tasks in many domains, such as tasks in e-commerce scenarios.
Approach: They propose a knowledge-injected pre-trained language model that can be transferred to both natural language understanding and generation tasks.
Outcome: The proposed model significantly outperforms baselines across the board in e-commerce scenarios.
Summary Level Training of Sentence Rewriting for Abstractive Summarization (D19-54)

Copied to clipboard

Challenge: Existing models rely on sentence-level rewards or suboptimal labels to achieve summary-level ROUGE scores.
Approach: They propose a model that extracts salient sentences from a document and paraphrases them to generate a summary.
Outcome: The proposed model improves on CNN/Daily Mail and New York Times datasets.
AutoNLU: An On-demand Cloud-based Natural Language Understanding System for Enterprises (2020.aacl-demo)

Copied to clipboard

Challenge: AutoNLU is an on-demand cloud-based system that enables users to create and edit datasets and train and test different state-of-the-art NLU models.
Approach: They introduce an on-demand cloud-based system that provides an easy-to-use interface . they build powerful keyphrase extraction models that achieve state-of-the-art results .
Outcome: The proposed model achieves state-of-the-art on two public benchmarks and is easy to use and use.
A Scalable Neural Shortlisting-Reranking Approach for Large-Scale Domain Classification in Natural Language Understanding (N18-3)

Copied to clipboard

Challenge: Existing approaches to classify a given utterance into domains are costly and time-consuming.
Approach: They propose a shortlisting-reranking neural model for large-scale domain classification for IPDAs . they use extensive experiments on 1,500 IPDA domains to test their effectiveness .
Outcome: The proposed model is tested on 1,500 IPDA domains.
Mandarinograd: A Chinese Collection of Winograd Schemas (2020.lrec-1)

Copied to clipboard

Challenge: Mandarinograd is a corpus of Winograd Schemas in Mandarin Chinese . WS are hard to collect and few datasets are publicly available .
Approach: They introduce a corpus of Winograd Schemas in Mandarin Chinese . they describe the difficulties faced when building the corpus and explain how they overcome the anomalies.
Outcome: The proposed corpus of Winograd Schemas in Mandarin Chinese is hard to build and resistant to statistical methods.
MathPrompter: Mathematical Reasoning using Large Language Models (2023.acl-industry)

Copied to clipboard

Challenge: Recent advances in natural language processing (NLP) can be attributed to massive scaling of Large Language Models (LLMs).
Approach: They propose a technique that improves performance of Large Language Models (LLMs) on arithmetic problems along with increased reliance in the predictions.
Outcome: The proposed technique improves performance on arithmetic problems and increases confidence in the output results.
When Choosing Plausible Alternatives, Clever Hans can be Clever (D19-60)

Copied to clipboard

Challenge: Pretrained language models have shown large improvements in the commonsense reasoning benchmark COPA, but recent work has identified superficial cues in benchmark datasets which are predictive of the correct answer.
Approach: They propose an extension of COPA that does not suffer from easy-to-exploit single token cues and exploits them.
Outcome: The proposed extension of COPA does not suffer from easy-to-exploit single token cues.
CorefInst: Leveraging LLMs for Multilingual Coreference Resolution (2026.tacl-1)

Copied to clipboard

Challenge: Existing methods for CR are encoder-only, decoder-based and asynchronous models.
Approach: They propose a multilingual CR methodology which leverages decoder-only LLMs to handle overt and zero mentions.
Outcome: The proposed model outperforms the leading multilingual CR model by 2 percentage points across all languages in the CorefUD v1.2 dataset.
Deep Bayesian Learning and Understanding (C18-3)

Copied to clipboard

Challenge: COLING 2018 is a conference for researchers and practitioners working on machine learning and deep learning.
Approach: a tutorial on machine learning and deep learning will be presented at COLING 2018 . the tutorial will focus on statistical models, deep neural networks, sequential learning and natural language understanding .
Outcome: This tutorial will present the latest advances in deep Bayesian and sequential learning at COLING 2018 .
Event Time Extraction and Propagation via Graph Attention Networks (2021.naacl-main)

Copied to clipboard

Challenge: Existing work on grounding events into a precise timeline has been limited due to the inherent ambiguity of language and the requirement for information propagation over inter-related events.
Approach: They propose a 4-tuple temporal representation for entity slot filling to ground events into a timeline using a graph attention network approach.
Outcome: The proposed approach yields 7.0% match rate over contextualized embedding approaches and 16.3% higher match rate compared to sentence-level manual event time argument annotation.
Don’t Shoot The Breeze: Topic Continuity Model Using Nonlinear Naive Bayes With Attention (2024.emnlp-industry)

Copied to clipboard

Challenge: Large-scale language models (LLMs) are becoming increasingly popular in business scenarios, but maintaining topic continuity is a challenge.
Approach: They propose a topic continuity model that assesses whether a response aligns with the initial conversation topic using a Naive Bayes approach.
Outcome: The proposed model outperforms existing models in handling lengthy and complex conversations.
Enhancing Self-Attention with Knowledge-Assisted Attention Maps (2022.naacl-main)

Copied to clipboard

Challenge: Existing works of knowledge infusion depend on multi-task learning frameworks, which are inefficient and require large-scale retraining when new knowledge is considered.
Approach: They propose a method which integrates knowledge-generated attention maps into the self-attention mechanism and integrates it into the model.
Outcome: The proposed model outperforms existing methods on academic datasets and industry-scale ad relevance applications.
An adaptable task-oriented dialog system for stand-alone embedded devices (P19-3)

Copied to clipboard

Challenge: a proposed speech-based task-oriented dialogue system is built on a small embedded device . the system does not require internet connectivity because all components run locally on the device - a cost-effective solution .
Approach: They propose a spoken-language end-to-end task-oriented dialogue system for small embedded devices such as home appliances.
Outcome: The proposed system is based on a demo run offline on swiss raspberry pi . it eliminates privacy risks and eliminates server costs and latency .
Dyna-bAbI: unlocking bAbI’s potential with dynamic synthetic benchmarking (2022.starsem-1)

Copied to clipboard

Challenge: Controlled synthetic tasks are an important resource for diagnosing model behavior.
Approach: They propose a framework that provides fine-grained control over task generation in bAbI.
Outcome: The proposed framework provides fine-grained control over task generation in the bAbI benchmark.
Cross-Lingual Dialogue Dataset Creation via Outline-Based Generation (2023.tacl-1)

Copied to clipboard

Challenge: Multilingual task-oriented dialogue (ToD) datasets suffer from severe limitations, such as being small in scale and lacking naturalness and cultural specificity in the target language.
Approach: They propose a novel outline-based annotation process where domain-specific abstract schemata of dialogue are mapped into natural language outlines.
Outcome: The proposed approach improves understanding, dialogue state tracking, and end-to-end dialogue evaluation in Arabic, Indonesian, Russian, and Kiswahili.
Delexicalized Paraphrase Generation (2020.coling-industry)

Copied to clipboard

Challenge: Using convolutional neural networks, we generate delexicalized sentences . 1.29% accuracy is achieved with the generated paraphrases .
Approach: They propose a neural paraphrasing model that generates delexicalized sentences . they use convolutional neural networks to pool on slot values and use pointers to locate them .
Outcome: The proposed model generates delexicalized sentences with high quality . it can be used for intent classification and named entity recognition tasks .
Active Learning for New Domains in Natural Language Understanding (N19-2)

Copied to clipboard

Challenge: Existing approaches to improve the accuracy of new domains are lacking annotated live utterances.
Approach: They propose an algorithm called Majority-CRF that uses an ensemble of classification models to guide the selection of relevant utterances and a sequence labeling model to prioritize informative examples.
Outcome: The proposed algorithm achieves 6.6%-9% error rate reduction and statistically significant improvements on six new domains.
Commonsense inference in human-robot communication (D19-60)

Copied to clipboard

Challenge: a gap exists in natural language understanding of commands between humans and machines.
Approach: They propose a method for commonsense inference to transform high-level commands into action commands for robotic systems to execute.
Outcome: The proposed method allows to build a knowledge base that consists of a large set of commonsense inferences.
The Importance of Being Parameters: An Intra-Distillation Method for Serious Gains (2022.emnlp-main)

Copied to clipboard

Challenge: Recent pruning methods remove redundant parameters according to parameter sensitivity, a gradient-based measure reflecting the contribution of the parameters.
Approach: They propose a general task-agnostic method to balance parameter sensitivity and a novel adaptive learning method to control strength of intra-distillation loss for faster convergence.
Outcome: The proposed method can reduce redundant parameters by over 80% without obvious performance degradation.
CogCompTime: A Tool for Understanding Time in Natural Language (D18-2)

Copied to clipboard

Challenge: Existing systems that extract temporal information from text can be useful for natural language understanding.
Approach: They propose a system that extracts temporal information from text and normalizes it to a standard format.
Outcome: The proposed system achieves state-of-the-art performance and incorporates the most recent progress.
Exploiting In-Domain Bilingual Corpora for Zero-Shot Transfer Learning in NLU of Intra-Sentential Code-Switching Chatbot Interactions (2022.emnlp-industry)

Copied to clipboard

Challenge: Multilingual speakers outnumber monolingual speakers in the world . CS is a frequent habit in both spoken and written informal communications .
Approach: They evaluate the efficacy of cross-lingual transfer learning with mBERT for NLU on a Basque-Spanish CS chatbot corpus.
Outcome: The proposed model outperforms models trained on Basque and Spanish without CS on a basque-Spanish chatbot corpus.
Sketching a Linguistically-Driven Reasoning Dialog Model for Social Talk (2022.acl-srw)

Copied to clipboard

Challenge: a new study shows that dialog systems that can hold social talk and make sense of conversational content are not efficient for context-sensitive natural language understanding and reasoning.
Approach: They propose a linguistically-informed architecture to handle social talk in English . they propose linguistic models that fit the context-sensitive components into a Bayesian game-theoretic model .
Outcome: The proposed architecture is based on corpus-based methods but does not track what is happening in a conversation.
Are the Tools up to the Task? an Evaluation of Commercial Dialog Tools in Developing Conversational Enterprise-grade Dialog Systems (N19-2)

Copied to clipboard

Challenge: Existing toolsets are incomplete in meeting the goal of building effective dialog systems, authors say .
Approach: They compare dialog tools available from a number of companies to determine their strengths and weaknesses . they provide quantitative and qualitative results in three main areas: natural language understanding, dialog, and text generation .
Outcome: The toolsets are incomplete, but they are compared to other tools to determine their strengths and weaknesses.
Infusing Finetuning with Semantic Dependencies (2021.tacl-1)

Copied to clipboard

Challenge: Several diagnostics help to localize the benefits of our approach.
Approach: They apply convolutional graph encoders to integrate semantic parses into task-specific finetuning.
Outcome: The proposed approach yields benefits to natural language understanding (NLU) tasks in the GLUE benchmark.
NarraSum: A Large-Scale Dataset for Abstractive Narrative Summarization (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing studies focus on summarizing news documents or structured documents.
Approach: They propose to use a large-scale narrative summarization dataset to encourage research . they find there is a performance gap between humans and the models on NarraSum .
Outcome: The proposed dataset shows that humans and state-of-the-art models perform poorly when summarizing a narrative . it contains 122K narratives collected from synopses of movies and TV episodes with diverse genres .
VLUE: A New Benchmark and Multi-task Knowledge Transfer Learning for Vietnamese Natural Language Understanding (2024.findings-naacl)

Copied to clipboard

Challenge: a lack of standard evaluation metrics and benchmarks makes it difficult to identify strengths of Vietnamese NLP models.
Approach: They propose to establish a standardized set of benchmarks for Vietnamese NLU . they propose to evaluate Vietnamese language understanding models using a pre-trained model .
Outcome: The proposed model combines proficiency of a multilingual pre-trained model with Vietnamese linguistic knowledge.
Alexa Conversations: An Extensible Data-driven Approach for Building Task-oriented Dialogue Systems (2021.naacl-demos)

Copied to clipboard

Challenge: Traditional goal-oriented dialogue systems require annotations which are hard to obtain for every new domain, limiting scalability.
Approach: They propose a data-driven approach to building goal-oriented dialogue systems . they use a seed dialogue simulator to generate annotated conversations instead of collecting annotations .
Outcome: The proposed system improves turn-level action signature prediction accuracy by 50% . the system is scalable, extensible and data efficient .
Your Pretrained Model Tells the Difficulty Itself: A Self-Adaptive Curriculum Learning Paradigm for Natural Language Understanding (2025.acl-srw)

Copied to clipboard

Challenge: Existing curriculum learning approaches rely on manually defined difficulty metrics which may not accurately reflect the model’s own perspective.
Approach: They propose a self-adaptive curriculum learning paradigm that prioritizes fine-tuning examples based on difficulty scores predicted by pre-trained language models (PLMs) they evaluate four datasets covering binary and multi-class classification tasks.
Outcome: The proposed model leads to faster convergence and improved performance compared to standard random sampling.
A Novel Cartography-Based Curriculum Learning Method Applied on RoNLI: The First Romanian Natural Language Inference Corpus (2024.acl-long)

Copied to clipboard

Challenge: Natural language inference (NLI) is an actively studied topic serving as a proxy for natural language understanding.
Approach: They propose to use a Romanian NLI corpus to analyze sentence pairs . they use multiple machine learning methods to establish competitive baselines .
Outcome: The proposed model improves on the best model by employing a new curriculum learning strategy based on data cartography.
Efficient Long-Text Understanding with Short-Text Models (2023.tacl-1)

Copied to clipboard

Challenge: Existing transformer-based pretrained language models cannot be applied to long sequences due to their quadratic complexity.
Approach: They propose a simple approach to long sequences that re-uses battle-tested short-text pretrained LMs.
Outcome: The proposed approach is competitive with specialized models that are up to 50x larger and require a dedicated and expensive pretraining step.
Graph-Based Semi-Supervised Learning for Natural Language Understanding (D19-53)

Copied to clipboard

Challenge: Semi-supervised learning is an efficient method to augment training data from unlabeled data.
Approach: They propose semi-supervised learning models and their inductive variants for NLU and use them to find similar utterances and construct a graph.
Outcome: The proposed model improves the error rate of the model by 5% using publicly available NLU data and models.
KERMIT: Complementing Transformer Architectures with Encoders of Explicit Syntactic Interpretations (2020.emnlp-main)

Copied to clipboard

Challenge: Syntactic parsers are losing their centrality in downstream tasks due to the success of large-scale textual representation learners.
Approach: They propose to embed symbolic syntactic parse trees into artificial neural networks to visualize how syntax is used in inference.
Outcome: The proposed encoder can visualize how syntax is used in inference.
CRUISE: Cold-Start New Skill Development via Iterative Utterance Generation (P18-4)

Copied to clipboard

Challenge: Existing systems require developers to manually generate and annotate a large number of utterances.
Approach: They propose a system that guides ordinary software developers to build a high quality NLU engine from scratch.
Outcome: The proposed system shows that iterative pruning of incorrect utterances reduces human workload and cognitive load.
Leveraging Large Language Models for Conversational Multi-Doc Question Answering: The First Place of WSDM Cup 2024 (2025.findings-acl)

Copied to clipboard

Challenge: WSDM Cup 2024 presents a challenge for conversational multi-doc question answering using large language models . a hybrid training strategy is developed to make the most of in-domain unlabeled data .
Approach: They propose a conversational multi-doc question answering challenge in WSDM Cup 2024 . they adapt LLMs to the task, then devise a hybrid training strategy to make the most of unlabeled data.
Outcome: The proposed approach ranked 1st in the WSDM Cup 2024 challenge . it exploits the superior natural language understanding and generation capability of Large Language Models .
Relative Importance in Sentence Processing (2021.acl-short)

Copied to clipboard

Challenge: In natural language processing, the relative importance of words is usually interpreted with respect to a specific task.
Approach: They compare the relative importance of words in English language processing by humans and neural language models by using saliency methods.
Outcome: The proposed method could be used to interpret neural language models.
Adaptive Natural Language Generation for Task-oriented Dialogue via Reinforcement Learning (2022.coling-1)

Copied to clipboard

Challenge: In task-oriented dialogue systems, the role of the natural language generation component is to convert a system's intentions, called dialogue acts (DAs), into natural language utterances and to convey DAs accurately to users.
Approach: They propose a method for Adaptive Natural language generation for Task-Oriented dialogue via Reinforcement learning that incorporates a natural language understanding module into the objective function of RL.
Outcome: The proposed method generates adaptive utterances against speech recognition errors and the different vocabulary levels of users in a multi-world task-oriented dialogue system.
How Does Data Corruption Affect Natural Language Understanding Models? A Study on GLUE datasets (2022.starsem-1)

Copied to clipboard

Challenge: Existing studies on the performance of pre-trained language models on natural language understanding tasks have focused on the natural language inference and textual entailment tasks.
Approach: They propose to use corrupted data to fine-tune pre-trained language models to assess their language understanding capabilities.
Outcome: The proposed transformations can be applied to all but one NLU task and show that understanding the meaning of utterances is not required for high performance.
Comprehensive Multi-Dataset Evaluation of Reading Comprehension (D19-58)

Copied to clipboard

Challenge: Recent research aims to facilitate training and evaluation on several reading comprehension datasets at the same time.
Approach: They propose an evaluation server that reports performance on seven diverse reading comprehension datasets and includes synthetic augmentations to test models' ability to handle out-of-domain questions.
Outcome: The evaluation server performs on seven reading comprehension datasets, and collects and includes synthetic augmentations for these datasets to test models' ability to handle out-of-domain questions.
Design Challenges for a Multi-Perspective Search Engine (2022.findings-naacl)

Copied to clipboard

Challenge: a document retrieval system fails to deliver diverse and direct responses to controversial questions . classical document retrievals provide a ranked list of references to relevant but not necessarily trustworthy web documents .
Approach: They propose a perspective-oriented document retrieval paradigm to address these challenges . they propose sponses with different perspectives within topically-related web documents .
Outcome: The proposed system is based on a user survey and a prototype . it will be used to assess the utility and understanding of the system .
PerspectroScope: A Window to the World of Diverse Perspectives (P19-3)

Copied to clipboard

Challenge: PerspectroScope is a web-based system that lets users query a discussion-worthy natural language claim .
Approach: They propose a web-based system which lets users query a discussion-worthy natural language claim and extract and visualize various perspectives in support or against the claim.
Outcome: The proposed system lets users query a discussion-worthy natural language claim and extract and visualize various perspectives in support or against the claim.
EXPLORER: Exploration-guided Reasoning for Textual Reinforcement Learning (2024.eacl-long)

Copied to clipboard

Challenge: Text-based games (TBGs) combine natural language understanding with reasoning.
Approach: They propose an exploration-guided reasoning agent for textual reinforcement learning that integrates natural language with reasoning.
Outcome: The proposed agent outperforms baseline agents on TWG and TWC games.
ReAct Meets Industrial IoT: Language Agents for Data Access (2025.emnlp-industry)

Copied to clipboard

Challenge: a framework for domain-specific language agents is being developed for industrial automation . a novel approach to adapting these systems to domain-based applications poses new challenges .
Approach: They propose a framework for deploying domain-specific language agents that can query industrial sensor data using natural language.
Outcome: The proposed framework outperforms standard prompting baselines across multiple LLMs including smaller models.
Probabilistically Masked Language Model Capable of Autoregressive Generation in Arbitrary Word Order (2020.acl-main)

Copied to clipboard

Challenge: Large-scale pretrained language models such as masked language model (MLM) have brought significant improvements to many NLU and NLG tasks.
Approach: They propose a probabilistic masking scheme for the masked language model and a model with a uniform prior distribution on the masking ratio.
Outcome: The proposed model outperforms BERT on a bunch of downstream NLG tasks.
Lightweight Transformers for Conversational AI (2022.naacl-industry)

Copied to clipboard

Challenge: Commercial dialogue systems typically require a small footprint and fast execution time, but recent trends are in the other direction, resulting in difficulties in model deployment.
Approach: They build Transformer-based Language Models from scratch on large corpora of conversational data and compare their performance against BERT and other strong baselines on dialogue probing tasks.
Outcome: The proposed model outperforms existing models on dialogue probing tasks and can be fine-tuned on a single consumer GPU card.
BelarusianGLUE: Towards a Natural Language Understanding Benchmark for Belarusian (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in NLP, such as large language models, have had groundbreaking impact on the field.
Approach: They propose a benchmark for Belarusian, an East Slavic language, with 15K instances in five tasks: sentiment analysis, linguistic acceptability, word in context, Winograd schema challenge, textual entailment.
Outcome: The proposed model underperforms on sentiment analysis, linguistic acceptability, word in context, Winograd schema challenge and textual entailment, but is competitive for linguistic acceptance.
GLM: General Language Model Pretraining with Autoregressive Blank Infilling (2022.acl-long)

Copied to clipboard

Challenge: Existing pretraining frameworks do not perform well for all tasks of three main categories, such as natural language understanding (NLU), unconditional generation, and conditional generation.
Approach: They propose a general language model based on autoregressive blank infilling to address this challenge.
Outcome: The proposed model outperforms BERT, T5, and GPT on a wide range of tasks across NLU, conditional and unconditional generation tasks.
Extracting Temporal Event Relation with Syntax-guided Graph Transformer (2022.findings-naacl)

Copied to clipboard

Challenge: Temporal relationship extraction is crucial for understanding complex events and reasoning over them.
Approach: They propose a Syntax-guided Graph Transformer network to extract temporal relations between events by explicitly exploiting the connection between two events based on their dependency parsing trees.
Outcome: The proposed approach outperforms state-of-the-art methods on MATRES and TB-DENSE with up to 7.9% absolute F-score gain.
LoRA-MGPO: Mitigating Double Descent in Low-Rank Adaptation via Momentum-Guided Perturbation Optimization (2025.findings-emnlp)

Copied to clipboard

Challenge: Low-Rank Adaptation (LoRA) adapts large language models by training only a small fraction of parameters, but as the rank of the low-rank matrices increases, LoRA exhibits an unstable “double descent” phenomenon, which delays convergence and impairs generalization by causing instability due to the attraction to sharp local minima.
Approach: They propose a framework that incorporates Momentum-Guided Perturbation Optimization (MGPO) MGPO stabilizes training dynamics by mitigating double descent phenomenon and guiding weight perturbations using momentum vectors from the optimizer’s state.
Outcome: The proposed framework improves performance on natural language understanding benchmarks and shows that it improves convergence and generalization.
A Large Resource of Patterns for Verbal Paraphrases (L18-1)

Copied to clipboard

Challenge: Xu et al., 2015: paraphrases play an important role in natural language understanding . he says it is difficult to propose a paraphrasing relation for natural language processing systems .
Approach: They propose a resource of such paraphrases that can be used to identify hidden paraphrase pairs . they propose to use the resource to identify paraphrase relationships between two words .
Outcome: The proposed resource contains tens of thousands of such pairs and is available for academic purposes.
Data Selection for Fine-tuning Large Language Models Using Transferred Shapley Values (2023.acl-srw)

Copied to clipboard

Challenge: Large language models (LMs) have been shown to be highly effective for identifying harmful training instances, but dataset size and model complexity constraints limit the ability to apply Shapley-based data valuation to fine-tuning large pre-trained language models.
Approach: They propose an algorithm that aggregates Shapley values from subsets for valuation of entire training set and a value transfer method that leverages value information extracted from a simple classifier trained using representations from the target language model.
Outcome: The proposed method outperforms existing methods on benchmark datasets and can filter fine-tuning data to increase language model performance compared to training with the full fine-uning dataset.
ArgGen: Prompting Text Generation Models for Document-Level Event-Argument Aggregation (2022.findings-aacl)

Copied to clipboard

Challenge: Existing discourse-level information extraction tasks are extractive in nature, but extracting information from larger bodies of discourse-like documents requires more natural language understanding and reasoning capabilities.
Approach: They propose a conditional text generation approach which generates consolidated event-arguments at a document-level with minimal loss of information.
Outcome: The proposed approach generates document-level argument spans in a low-resource and zero-shot setting and can be leveraged in other related multilingual text generation tasks.
AMBERT: A Pre-trained Language Model with Multi-Grained Tokenization (2021.findings-acl)

Copied to clipboard

Challenge: Pre-trained language models such as BERT have shown great power in natural language understanding . fine-grained tokenizations have advantages and disadvantages for learning of pre-tried models .
Approach: They propose a pretrained language model based on both fine-grained and coarse-grain tokenizations . they propose to use both tokenization techniques to learn pre-trained models .
Outcome: The proposed model outperforms BERT on benchmark datasets for Chinese and English . it can perform better with the same computational cost as BERT, the authors show .
KorNLI and KorSTS: New Benchmark Datasets for Korean Natural Language Understanding (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmark datasets for natural language inference and semantic textual similarity (STS) are not available in the Korean language.
Approach: They construct and release new datasets for Korean NLI and STS . they machine-translate existing English training sets and manually translate development and test sets into Korean to accelerate research on Korean NLU.
Outcome: The proposed datasets are available at https://github.com/kakaobrain/KorNLUDatasets.
Snoopy: An Online Interface for Exploring the Effect of Pretraining Term Frequencies on Few-Shot LM Performance (2022.emnlp-demos)

Copied to clipboard

Challenge: Snoopy allows researchers to analyze the impact of the overlap between pretraining corpus and test data on model performance statistics.
Approach: They propose to align terms in test instances with their frequency in the Pile to explore correlations between model accuracy and model size and number.
Outcome: Snoopy allows researchers to analyze term frequency statistics in large language models on NLP benchmarks.
TableFormer: Robust Transformer Modeling for Table-Text Encoding (2022.acl-long)

Copied to clipboard

Challenge: Existing tables models require linearization of the table structure, where row or column order is encoded as an unwanted bias.
Approach: They propose a robust and structurally aware table-text encoding architecture TableFormer where tabular structural biases are incorporated completely through learnable attention biase.
Outcome: The proposed architecture outperforms strong baselines on SQA, WTQ and TabFact table reasoning datasets and achieves state-of-the-art performance on SQ.
To Chat or Task: a Multi-turn Dialogue Generation Framework for Task-Oriented Dialogue Systems (2025.acl-industry)

Copied to clipboard

Challenge: Large language models (LLMs) are designed to handle complex task requests, but lack of specific datasets for training and evaluation of such systems .
Approach: They propose a framework to generate a dataset for in-vehicle speech recognition systems . they train an in-car context sensor that correctly identifies the functional intent of the driver .
Outcome: The proposed framework outperforms baseline models across experimental settings.
PromptDA: Label-guided Data Augmentation for Prompt-based Few Shot Learners (2023.eacl-main)

Copied to clipboard

Challenge: Existing studies on prompt-based few-shot tuning focus on deriving proper label words with a verbalizer or generating prompt templates to elicit semantics from PLMs.
Approach: They propose a framework that leverages label semantics for prompt-based tuning.
Outcome: The proposed framework improves on few-shot text classification tasks by leveraging label semantics and data augmentation.
Recent Neural Methods on Slot Filling and Intent Classification for Task-Oriented Dialogue Systems: A Survey (2020.coling-main)

Copied to clipboard

Challenge: In recent years, neural-network based models have been used for a wide range of tasks, including slot filling and intent classification.
Approach: They propose three neural architectures to model slot filling and intent classification . they propose independent models, joint models and transfer learning models that exploit the mutual benefit of the two tasks simultaneously and scale the model to new domains.
Outcome: The proposed models model SF and IC separately, exploit mutual benefit of the two tasks simultaneously and scale the model to new domains.
The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants (2024.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for text comprehension only cover 30 languages, but lack of labeled data is a major obstacle to building functional systems in most languages.
Approach: They present a multiple-choice machine reading comprehension dataset spanning 122 languages . they use it to evaluate the capabilities of multilingual masked language models and large language models .
Outcome: The proposed dataset enables the evaluation of text models in high-, medium- and low-resource languages.
Text-based NP Enrichment (2022.tacl-1)

Copied to clipboard

Challenge: Existing NLP tasks and benchmarks do not cover all NP-mediated relations . we aim to enrich each NP in a text with all the preposition-mediated relationships that hold between it and other NPs in the text.
Approach: They propose a task to enrich NPs with preposition-mediated relations that hold between them . they build a large-scale dataset and analyze the data to test the task .
Outcome: The proposed task is based on a large-scale dataset and fine-tuned language models.
Towards General-Domain Word Sense Disambiguation: Distilling Large Language Model into Compact Disambiguator (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for Word Sense Disambiguation rely heavily on manually annotated data, which limits coverage and generalization.
Approach: They propose a framework that leverages large language models as knowledge distillers to build silver-standard WSD corpora by combining generation-based distillation and annotation-based disambiguation.
Outcome: The proposed framework outperforms existing methods on general-domain benchmarks by 50% on the most challenging test set and by 1000 times fewer parameters.
SpeechLLMs for Large-scale Contextualized Zero-shot Slot Filling (2025.emnlp-industry)

Copied to clipboard

Challenge: Slot filling is a key subtask in spoken language understanding (SLU) . recent advent of speech-based large language models has opened new avenues for speech understanding .
Approach: They propose to improve slot-filling task by creating an empirical upper bound for the task . they propose to use a speech-based large language model to integrate speech and text modalities .
Outcome: The proposed model improves slot filling performance while reducing generalization gaps.
Inducer-tuning: Connecting Prefix-tuning and Adapter-tuning (2022.emnlp-main)

Copied to clipboard

Challenge: Prefix-tuning is an essential paradigm of parameter-efficient transfer learning . fine-tuned models require separate copies of model parameters for each task .
Approach: They propose to understand and further develop prefix-tuning through the kernel lens . they propose a new variant of prefix tuning that shares the exact mechanism as prefix tun .
Outcome: The proposed method improves prefix-tuning performance by training only a small portion of parameters.
Syntactic Structure Distillation Pretraining for Bidirectional Encoders (2020.tacl-1)

Copied to clipboard

Challenge: Textual representation learners trained on large amounts of data have been successful on downstream tasks.
Approach: They propose a knowledge distillation strategy for injecting syntactic biases into BERT pretraining by distilling the approximate marginal distribution over words in context from the syntaktic LM.
Outcome: The proposed method reduces relative error by 2–21% on a diverse set of structured prediction tasks.
Linking artificial and human neural representations of language (D19-1)

Copied to clipboard

Challenge: a pre-trained BERT architecture is used to fine-tune sentence encoding models on a variety of natural language understanding (NLU) tasks.
Approach: They compare sentence encoding models with fMRI-based fMR predictions of the sentence . they use a pre-trained BERT architecture as a baseline and fine-tune it on a variety of natural language understanding (NLU) tasks.
Outcome: The proposed model does not yield significant improvements in brain decoding performance on the natural language understanding (NLU) tasks.
Unsupervised Neologism Normalization Using Embedding Space Mapping (D19-55)

Copied to clipboard

Challenge: Neologisms refer to recent expressions that are specific to certain entities or events, but have not yet been accepted into mainstream language.
Approach: They propose an unsupervised approach for detecting and normalizing neologisms in social media content without relying on parallel training data.
Outcome: The proposed method detects neologisms and normalizes them to canonical words without training data.
Benchmarking Robustness of Machine Reading Comprehension Models (2021.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks only evaluate models' robustness under test-time perturbations or adversarial attacks.
Approach: They propose a model-agnostic benchmark to evaluate models' robustness under adversarial attacks.
Outcome: The proposed model-agnostic benchmark evaluates models under four different types of adversarial attacks.
Plug and Play Knowledge Distillation for kNN-LM with External Logits (2022.aacl-short)

Copied to clipboard

Challenge: Despite the promising evaluation results by knowledge distillation (KD) in natural language understanding (NLU) and sequence-to-sequence (seq2sequ) tasks, KD for causal language modeling (LM) remains a challenge.
Approach: They propose to use external logits to improve a student's kNN-LM by leveraging teacher's knowledge at test time.
Outcome: The proposed method improves a student's kNN-LM in multiple language modeling datasets and improves perplexity.
Scalable Prompt Generation for Semi-supervised Learning with Language Models (2023.findings-eacl)

Copied to clipboard

Challenge: Prompt-based learning methods in semi-supervised learning (SSL) settings have been shown to be effective on multiple natural language understanding datasets and tasks.
Approach: They propose to use a set of prompt tokens to create diverse prompt models and a varying number of soft prompt token to encourage language models to learn different prompts.
Outcome: The proposed method achieves the best average accuracy of 71.5% in different few-shot learning settings.
VIRT: Improving Representation-based Text Matching via Virtual Interaction (2022.emnlp-main)

Copied to clipboard

Challenge: Experimental results show that representation-based text matching methods suffer from performance degradation due to the lack of interactions between the pair of texts.
Approach: They propose a virtual interaction mechanism that enables deep interaction between texts . they propose 'inteRacTion mechanism' that can be integrated into existing methods as plugins .
Outcome: The proposed method outperforms state-of-the-art models on six text matching benchmarks.
Probing the Depths of Language Models’ Contact-Center Knowledge for Quality Assurance (2024.emnlp-industry)

Copied to clipboard

Challenge: Recent advances in large Language Models (LMs) have significantly enhanced their capabilities across various domains, including natural language understanding and domain knowledge.
Approach: They propose methods to transfer domain-specific knowledge to smaller models by leveraging evaluation plans generated by more knowledgeable models with optional human-in-the-loop refinement to enhance the capabilities of smaller models.
Outcome: The proposed models improve 18.95% on an in-house QA dataset on a contact-center quality assurance task.
Structural Persistence in Language Models: Priming as a Window into Abstract Language Representations (2022.tacl-1)

Copied to clipboard

Challenge: a rich literature has emerged in the last few years addressing these questions, including whether specific LMs have acquired specific linguistic constructions.
Approach: They introduce a novel metric and release Prime-LM, a large corpus where they control for various linguistic factors that interact with priming strength.
Outcome: The proposed model can learn abstract structural information independent of the structure of a sentence and is able to perform tasks that require natural language understanding skills.
GlyphPattern: An Abstract Pattern Recognition for Vision-Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for abstract pattern recognition are easier because they do not involve a natural language description of the pattern.
Approach: They present a dataset that pairs human-written descriptions of visual patterns with three visual presentation styles.
Outcome: The proposed benchmark pairs human-written and human-verified patterns with three visual presentation styles.
Towards Unsupervised Language Understanding and Generation by Joint Dual Learning (2020.acl-main)

Copied to clipboard

Challenge: Existing work exploits dual property between understanding and generation to improve performance of modular dialogue systems.
Approach: They propose a dual supervised learning framework that exploits the dual property between understanding and generation.
Outcome: The proposed framework improves both NLU and NLG performance by incorporating supervised and unsupervised learning algorithms.
WALNUT: A Benchmark on Semi-weakly Supervised Learning for Natural Language Understanding (2022.naacl-main)

Copied to clipboard

Challenge: Existing studies on weak supervision for NLU focus on a specific task or simulate weak supervision signals from ground-truth labels.
Approach: They propose a benchmark to advocate and facilitate research on weak supervision for NLU . they use document-level and token-level prediction tasks as examples .
Outcome: The proposed benchmark advocates and facilitates research on weak supervision for NLU tasks.
Pragmatic Perspective on Assessing Implicit Meaning Interpretation in Sentiment Analysis Models (2025.acl-srw)

Copied to clipboard

Challenge: Using pragmatic theories of implicature, interpreting texts with implicit meaning correctly is essential for precise natural language understanding.
Approach: They propose to use transformer models fine-tuned for sentiment analysis to illustrate the challenges in computational interpretation of implicatures.
Outcome: The proposed model classifications reveal the limitations of supervised machine learning methods in detecting implicit sentiments.
FedEx-LoRA: Exact Aggregation for Federated and Efficient Fine-Tuning of Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for low-rank averaging of LoRA adapters result in inexact updates.
Approach: They propose a method which adds a residual error term to the pre-trained frozen weight matrix to achieve exact updates with minimal computational and communication overhead.
Outcome: The proposed method achieves exact updates with minimal computational and communication overhead, preserving LoRA’s efficiency.
RiSAWOZ: A Large-Scale Multi-Domain Wizard-of-Oz Dataset with Rich Semantic Annotations for Task-Oriented Dialogue Modeling (2020.emnlp-main)

Copied to clipboard

Challenge: RiSAWOZ contains 11.2K human-to-human (H2H) multi-turn semantically annotated dialogues spanning over 12 domains . despite of substantial progress made, there are challenges in creating challenging datasets in terms of size, multiple domains, semantic annotations and complexity.
Approach: They propose a large-scale multi-domain Chinese Wizard-of-Oz dataset with rich semantic annotations that captures discourse phenomena for task-oriented dialogue modeling.
Outcome: The proposed dataset contains 11.2K human-to-human (H2H) multi-turn semantically annotated dialogues with more than 150K utterances spanning over 12 domains.
ParsiNLU: A Suite of Language Understanding Challenges for Persian (2021.tacl-1)

Copied to clipboard

Challenge: Despite progress in natural language understanding, most progress is concentrated on resource-rich languages like English . despite high-quality benchmarks, there are few available NLU datasets for Persian language .
Approach: They propose a benchmark for Persian language that includes a range of language understanding tasks . they present their results on monolingual and multilingual pre-trained language models .
Outcome: The proposed benchmarks compare human performance with monolingual and multilingual models on Persian language with high quality evaluation datasets.
Improving Commonsense Contingent Reasoning by Pseudo-data and Its Application to the Related Tasks (2022.coling-1)

Copied to clipboard

Challenge: Contingent reasoning is one of the essential abilities in natural language understanding . despite advances in deep learning, the task of contingent reasoning is still difficult for computers .
Approach: They propose to generate large-scale pseudo-problems and incorporate them into training . they also investigate the generality of contingent knowledge through quantitative evaluation .
Outcome: The proposed method is able to evaluate the generality of contingent knowledge through transfer learning.
On Curriculum Learning for Commonsense Reasoning (2022.naacl-main)

Copied to clipboard

Challenge: Recent research suggests that data order can have a significant impact on the performance of finetuned models for natural language understanding.
Approach: They use paced curriculum learning to rank data and sample training mini-batches with increasing levels of difficulty during finetuning.
Outcome: The proposed model improves performance for socialIQA, CosmosQA, CODAH, HellaSwag, WinoGrande in both tuning settings.
Enhancing Input-Label Mapping in In-Context Learning with Contrastive Decoding (2025.acl-short)

Copied to clipboard

Challenge: Prior research has found that large language models overlook input-label mapping information in ICL, relying more on their pre-trained knowledge.
Approach: They propose a novel method that contrasts input-label mappings between positive and negative in-context examples to improve model performance.
Outcome: The proposed method improves performance on 7 natural language understanding tasks without additional training.
Data-Centric Perspectives on Agentic Retrieval-Augmented Generation: A Survey (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) excel at natural language understanding and generation, yet rely on static pre-training data.
Approach: They propose to augment Large Language Models with external retrieval to ground model outputs . traditional RAG is constrained by a fixed retrieve-then-generate routine . authors aim to guide creation of high-quality datasets for next generation of adaptive LLM agents .
Outcome: The proposed model can decompose tasks, issue exploratory queries, and refine evidence through iterative retrieval.
Multi-Task Networks with Universe, Group, and Task Feature Learning (P19-1)

Copied to clipboard

Challenge: In multi-task learning, multiple related tasks are learned together.
Approach: They propose methods that take advantage of natural groupings of related tasks . they propose parallel and serial architectures that can learn different feature spaces .
Outcome: The proposed methods improve performance on natural language understanding (NLU) tasks.
Task-adaptive Pre-training and Self-training are Complementary for Natural Language Understanding (2021.findings-emnlp)

Copied to clipboard

Challenge: Task-adaptive pre-training (TAPT) and Self-training can be complementary with simple TFS protocol.
Approach: They propose to use task-adaptive pre-training and self-training to combine TAPT and ST with a simple TFS protocol to achieve strong combined gains across six datasets.
Outcome: The proposed method can achieve strong combined gains across six datasets covering sentiment classification, paraphrase identification, natural language inference, named entity recognition and dialogue slot classification.
Investigating the (De)Composition Capabilities of Large Language Models in Natural-to-Formal Language Conversion (2025.naacl-long)

Copied to clipboard

Challenge: Existing frameworks for evaluating the decomposition and composition capabilities of large language models (LLMs) in N2F are inadequate, and there are errors that can be attributed to deficiencies in natural language understanding and the learning and use of symbolic systems.
Approach: They propose a framework that semi-automatically performs sample and task construction . main findings include that LLMs are deficient in both decomposition and composition .
Outcome: The proposed framework evaluates the most advanced LLMs on a variety of common formal languages.
CoDA21: Evaluating Language Understanding Capabilities of NLP Models With Context-Definition Alignment (2022.acl-short)

Copied to clipboard

Challenge: Pretrained language models (PLMs) have achieved superhuman performance on many benchmarks, creating a need for harder tasks.
Approach: They propose a benchmark that measures natural language understanding (NLU) abilities of pretrained language models.
Outcome: The proposed benchmark measures the ability of pretrained language models to perform on many tasks.
Unsupervised Deep Structured Semantic Models for Commonsense Reasoning (N19-1)

Copied to clipboard

Challenge: Existing methods for commonsense reasoning rely on human-crafted features and knowledge bases, but unsupervised learning is not feasible due to the lack of labeled training data or comprehensive knowledge bases.
Approach: They propose two unsupervised models based on the Deep Structured Semantic Models framework to tackle two commonsense reasoning tasks: Winograd Schema Challenge (WSC) and Pronoun Disambiguation (PDP).
Outcome: The proposed models capture contextual information in the sentence and co-reference information between pronouns and nouns, and achieve significant improvement over previous state-of-the-art approaches.
Modelling Language Acquisition through Syntactico-Semantic Pattern Finding (2023.findings-eacl)

Copied to clipboard

Challenge: Usage-based theories of language acquisition have documented the processes by which children acquire language through communicative interaction.
Approach: They propose a method for learning grammars based on similarities and differences in linguistic observations alone.
Outcome: The proposed method is able to learn compositional lexical and item-based constructions of variable extent and degree of abstraction, along with a network of emergent syntactic categories.
Language Models are Crossword Solvers (2025.naacl-long)

Copied to clipboard

Challenge: Modern crossword models demonstrate astounding skills in reasoning, coding, wordplay, question answering, and a multitude of other tasks.
Approach: They propose a search algorithm that generalizes well and can support answers with sound rationale by solving full crossword grids with out-of-the-box LLMs.
Outcome: The proposed model outperforms state-of-the-art models in solving crossword grids for the first time and generalizes well.
ConCodeEval: Evaluating Large Language Models for Code Constraints in Domain-Specific Languages (2025.acl-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated potential in code generation and natural language understanding, but they struggle with code constraints.
Approach: They propose to use Large Language Models to handle constraints represented in code . they use JSON, YAML, XML, Python, and natural language to test their effectiveness .
Outcome: The proposed benchmark shows that LLMs can handle code constraints better than natural language . the results suggest that conscious choice of representations can lead to optimal use of LLM in enterprise use cases involving code constraints.
Metacognitive Prompting Improves Understanding in Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Recent advances in prompting have enhanced reasoning in logic-intensive tasks for LLMs, yet the nuanced understanding abilities of these models remain underexplored.
Approach: They propose a strategy inspired by human introspective reasoning processes to enhance LLMs' understanding abilities.
Outcome: The proposed method outperforms chain-of-thought prompting and its advanced versions on ten natural language understanding (NLU) datasets.
Supervised Domain Enablement Attention for Personalized Domain Classification (D18-1)

Copied to clipboard

Challenge: Recent IPDAs cover more than several thousands of diverse domains including Alexa Skills, Google Actions, and Cortana Skills.
Approach: They propose a supervised enablement attention mechanism that utilizes sigmoid activation for the attention weighting and self-distillation to leverage the attention information of other enabled domains.
Outcome: The proposed approach improves domain classification performance on real-world domains.
Are We Modeling the Task or the Annotator? An Investigation of Annotator Bias in Natural Language Understanding Datasets (D19-1)

Copied to clipboard

Challenge: Having only a few workers generate the majority of dataset examples raises concerns about data diversity .
Approach: They perform a series of experiments to investigate annotator biases in recent NLU datasets . they find that models are able to recognize the most productive annotators .
Outcome: The results show that models can recognize the most productive annotators and do not generalize well to examples from annotator that did not contribute to the training set.
Investigating Meta-Learning Algorithms for Low-Resource Natural Language Understanding Tasks (D19-1)

Copied to clipboard

Challenge: Existing methods to learn general representations of text can achieve sub-optimal performance in low-resource scenarios.
Approach: They propose to use language model pre-training and multi-task learning to learn robust representations but these methods can achieve sub-optimal performance in low-resource scenarios.
Outcome: The proposed model outperforms strong baselines on the GLUE benchmark and can be adapted to new tasks efficiently and effectively.
The Impacts of Unanswerable Questions on the Robustness of Machine Reading Comprehension Models (2023.eacl-main)

Copied to clipboard

Challenge: Pretrained language models have achieved super-human performances on many Machine Reading Comprehension (MRC) benchmarks.
Approach: They propose to fine-tune three state-of-the-art language models on SQuAD 1.1 or SQu AD 2.0 and then evaluate their robustness under adversarial attacks.
Outcome: The proposed model is able to perform better under adversarial attacks than model fine-tuned on SQuAD 1.1 or 2.0.
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation (2025.findings-emnlp)

Copied to clipboard

Challenge: Prior work has shown that intent detection enhances LLMs’ moderation guardrails, but the robustness of these guardrail mechanisms under malicious manipulations remains under-explored.
Approach: They propose a two-stage intent-based prompt-refinement framework that first transforms harmful inquiries into structured outlines and further reframes them into declarative-style narratives.
Outcome: The proposed framework outperforms several cutting-edge jailbreak methods and evades even advanced Intent Analysis (IA) and Chain-of-Thought (CoT)-based defenses.
Temporal Relation Classification using Boolean Question Answering (2023.findings-acl)

Copied to clipboard

Challenge: a new approach for temporal relation classification (TRC) is proposed . a boolean question answering model is used to classify temporal relations between two events .
Approach: They propose an efficient approach for temporal relation classification using a boolean question answering model based on TRC annotation guidelines.
Outcome: The proposed model outperforms state-of-the-art models by 2.4% on questions designed by human annotation experts.
Debiasing Methods in Natural Language Understanding Make Bias More Accessible (2021.emnlp-main)

Copied to clipboard

Challenge: Recent debiasing methods in natural language understanding improve performance on out-of-distribution datasets by pressuring models into making unbiased predictions.
Approach: They propose a general probing-based framework that allows for post-hoc interpretation of biases in language models and use an information-theoretic approach to measure the extractability of certain biase .
Outcome: The proposed framework allows for post-hoc interpretation of biases in language models and measures the extractability of certain biase .
MoEBERT: from BERT to Mixture-of-Experts via Importance-Guided Adaptation (2022.naacl-main)

Copied to clipboard

Challenge: Existing methods for training pre-trained language models have limited practicality due to latency requirements.
Approach: They propose a method that uses a Mixture-of-Experts structure to increase model capacity and inference speed.
Outcome: The proposed method outperforms existing distillation methods on natural language understanding and question answering tasks.
Beyond Silent Letters: Amplifying LLMs in Emotion Recognition with Vocal Nuances (2025.findings-naacl)

Copied to clipboard

Challenge: Recent studies have demonstrated that Large Language Models possess a form of emotional intelligence, capable of interpreting emotional stimuli in text.
Approach: They propose a method that translates speech characteristics into natural language descriptions and integrates them into LLMs to perform multimodal emotion analysis via text prompts.
Outcome: The proposed method outperforms baseline models that require structural modifications on two datasets showing significant improvements in emotion recognition accuracy.
Modular Monolingual Adaptation using Pretrained Language Models (2026.acl-industry)

Copied to clipboard

Challenge: Existing approaches to building monolingual models for low-resource languages require a full model tuning process.
Approach: They propose a modular approach to build monolingual models for low-resource languages by finetuning the whole model on the target language.
Outcome: The proposed model improves on natural language understanding tasks on Scottish Gaelic, Irish, and Quechua with Quechuan being a very low-resource language.
D.Va: Validate Your Demonstration First Before You Use It (2025.acl-long)

Copied to clipboard

Challenge: In-context learning (ICL) heavily relies on selecting effective demonstrations to achieve outputs that better align with the expected results.
Approach: They propose a method which integrates a demonstration validation perspective into this field and integrates it into the learning paradigm.
Outcome: The proposed method surpasses all retrieval-based in-context learning techniques across both natural language understanding (NLU) and natural language generation (NLG) tasks.
RiddleSense: Reasoning about Riddle Questions Featuring Linguistic Creativity and Commonsense Knowledge (2021.findings-acl)

Copied to clipboard

Challenge: a riddle-style commonsense questions require complex commonsensense reasoning and figurative language skills . there is currently no dataset aimed at testing these abilities . authors propose a new multiple-choice question answering task .
Approach: They propose a new multiple-choice question answering task that uses a large dataset for riddlestyle commonsense questions.
Outcome: The proposed task comes with the first large dataset for answering riddlestyle commonsense questions.
mPMR: A Multilingual Pre-trained Machine Reader at Scale (2023.acl-short)

Copied to clipboard

Challenge: Existing mPLMs only transfer NLU capability from source to target languages . mPMR allows direct inheritance of multilingual NLU capabilities to downstream tasks .
Approach: They propose a method to guide multilingual pre-trained language models to perform natural language understanding in multiple languages.
Outcome: mPMR enables multilingual pre-trained language models to perform natural language understanding (NLU) in multiple languages.
mmT5: Modular Multilingual Pre-Training Solves Source Language Hallucinations (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent large language models display surprising multilingual capabilities despite being pre-trained on English data.
Approach: They propose a multilingual sequence-to-sequence model that disentangles language-specific information from language-agnostic information.
Outcome: The proposed model outperforms existing models on representative natural language understanding and generation tasks in 40+ languages.
Learning Slice-Aware Representations with Mixture of Attentions (2021.findings-acl)

Copied to clipboard

Challenge: Real-world machine learning systems are achieving excellent performance in terms of coarse-grained metrics like overall accuracy and F-1 score.
Approach: They extend slice-based learning (SBL) with a mixture of attentions to learn slice-aware dual attentive representations.
Outcome: The proposed approach outperforms the baseline method and the original SBL approach on monitored slices with two natural language understanding tasks.
A Sequence-to-Sequence Approach to Dialogue State Tracking (2021.acl-long)

Copied to clipboard

Challenge: Existing methods for dialogue state tracking are still challenging, but they are improving . a new approach to dialogue state monitoring is proposed, called Seq2Seq-DU .
Approach: They propose a new dialogue state tracking module that formalizes DST as a sequence-to-sequence problem.
Outcome: The proposed method outperforms existing methods on benchmark datasets in different settings.
Zero-Shot End-to-End Spoken Language Understanding via Cross-Modal Selective Self-Training (2024.eacl-long)

Copied to clipboard

Challenge: End-to-end (E2E) spoken language understanding models are constrained by the cost of collecting speech-semantics pairs.
Approach: They propose a model that learns E2E SLU without speech-semantics pairs . they propose cross-modal selective self-training (CMSST) to address imbalance and noise issues .
Outcome: The proposed model learns E2E SLU without speech-semantics pairs . the proposed model requires the domains of speech-text and text-sensitization to match .
Adversarial Robustness of Prompt-based Few-Shot Learning for Natural Language Understanding (2023.findings-acl)

Copied to clipboard

Challenge: Recent few-shot learning methods focus on improving downstream task performance, but there is limited understanding of the adversarial robustness of such methods.
Approach: They evaluate prompt-based FSL methods against fully fine-tuned models to better understand the impact of various factors towards robustness.
Outcome: The proposed methods show that they are less robust in the face of adversarial perturbations than fully fine-tuned models.
TESS: Text-to-Text Self-Conditioned Simplex Diffusion (2024.eacl-long)

Copied to clipboard

Challenge: Existing models for diffusion generation are expensive and discrete, resulting in a large number of diffusion steps to generate text.
Approach: They propose a text diffusion model that is fully non-autoregressive and employs a new form of self-conditioning and applies the diffusion process on the logit simplex space rather than the learned embedding space.
Outcome: The proposed model outperforms state-of-the-art non-autoregressive models, requires fewer diffusion steps with minimal drop in performance, and is competitive with pretrained autoregressive sequence-to-sequence models.
To What Extent Do Natural Language Understanding Datasets Correlate to Logical Reasoning? A Method for Diagnosing Logical Reasoning. (2022.coling-1)

Copied to clipboard

Challenge: Reasoning and knowledge-related skills are considered as fundamental skills for natural language understanding (NLU) tasks.
Approach: They propose a method to diagnose correlations between an NLU dataset and a specific skill.
Outcome: The proposed method is able to diagnose correlations between dataset and logical reasoning skill on 8 MRC and 3 NLI datasets.
Joint Energy-based Model Training for Better Calibrated Natural Language Understanding Models (2021.eacl-main)

Copied to clipboard

Challenge: Existing calibration methods rescale posterior distributions of classifiers after training.
Approach: They propose to use a noise contrastive estimation technique to train an energy-based model during finetuning of pretrained text encoders.
Outcome: The proposed model can reach a better calibration competitive to strong baselines with little or no loss in accuracy.
GLADIS: A General and Large Acronym Disambiguation Benchmark (2023.eacl-main)

Copied to clipboard

Challenge: Existing acronym disambiguation benchmarks are limited to specific domains . a study on a Microsoft question answering forum found that only 7% of acronyms co-occur with their corresponding long forms, which confuses the readers about the meaning of a text.
Approach: They propose a new acronym disambiguation benchmark with a dictionary and a pre-training corpus . they then pre-train a language model on the constructed corpus and show the challenges .
Outcome: The proposed benchmarks pre-train a language model on the constructed corpus for general acronym disambiguation.
DialogStudio: Towards Richest and Most Diverse Unified Dataset Collection for Conversational AI (2024.findings-eacl)

Copied to clipboard

Challenge: DialogStudio is the largest and most diverse collection of dialogue datasets . existing datasets lack diversity and comprehensiveness, authors say .
Approach: They introduce DialogStudio: the largest and most diverse collection of dialogue datasets . DialogStuio aggregates more than 80 diverse dialogue dataset .
Outcome: a new dataset is created to improve the quality and diversity of dialogue datasets . DialogStudio is the largest and most diverse collection of dialogue data .
Improving Constituency Parsing with Span Attention (2020.findings-emnlp)

Copied to clipboard

Challenge: Constituency parsing is a fundamental task for natural language understanding . n-grams are a conventional type of feature for contextual information . experimental results show that neural parsers with no grammar rules outperform statistical ones .
Approach: They propose to incorporate n-grams into span representations by weighting them according to their contributions to the parsing process.
Outcome: The proposed approach outperforms existing statistical grammar-based models on Arabic, Chinese, and English datasets.
A Neural-Symbolic Approach to Natural Language Understanding (2022.findings-emnlp)

Copied to clipboard

Challenge: Pre-trained language models have enabled deep neural networks to perform natural language understanding tasks, but their performance can drastically deteriorate when logical reasoning is needed.
Approach: They propose a framework for NLU based on analogical reasoning based upon neural processing and logical reasoning using both neural and symbolic processing.
Outcome: The proposed framework outperforms state-of-the-art methods on two NLU tasks, question answering (QA) and natural language inference (NLI).
Grounding learning of modifier dynamics: An application to color naming (D19-1)

Copied to clipboard

Challenge: Existing models for grounding are unable to understand modified color expressions, such as “light blue”.
Approach: They propose a model that learns more complex transformations in RGB space and a hard ensemble model that selects a color space depending on the modifier-color pair.
Outcome: The proposed model performs better in the HSV color space than the state-of-the-art model.
Question Modifiers in Visual Question Answering (2022.lrec-1)

Copied to clipboard

Challenge: Visual Question Answering (VQA) is a multi-disciplinary task that requires integration of several key disciplines.
Approach: They develop a model that adds modifiers to questions based on object properties and spatial relationships using Amazon Mechanical Turk data.
Outcome: The proposed model can improve when questions are modified to include more details.
Does Masked Language Model Pre-training with Artificial Data Improve Low-resource Neural Machine Translation? (2023.findings-eacl)

Copied to clipboard

Challenge: Pre-training masked language models with artificial data has been proven beneficial for several natural language processing tasks, however, it has been less explored for neural machine translation (NMT).
Approach: They pre-trained masked language models with random sequences and created artificial data mimicking token frequency information from the real world.
Outcome: The results show that pre-training models with artificial data improves translation performance in low-resource situations.
Effect of Visual Extensions on Natural Language Understanding in Vision-and-Language Models (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for creating vision-and-language models involve structural modifications and V&L pre-training.
Approach: They propose to extend a language model through structural modifications and V&L pre-training to make it inherit the capability of natural language understanding from the original language model.
Outcome: The proposed method improves performance of vision-and-language models by extending pre-trained models with the same pre-training.
AROMA: Autonomous Rank-one Matrix Adaptation (2025.emnlp-main)

Copied to clipboard

Challenge: Low-rank adaptation (LoRA) and adaptive low-rank adaption (AdaLoRa) are effective for large language models but are expensive as model sizes escalate into hundreds of billions of parameters.
Approach: They propose a framework that automatically builds up rank-one components with very few trainable parameters that gradually diminish to zero.
Outcome: The proposed framework significantly reduces parameters compared to LoRA and AdaLoRA while maintaining subspace independence.
A Dog Is Passing Over The Jet? A Text-Generation Dataset for Korean Commonsense Reasoning and Evaluation (2022.findings-naacl)

Copied to clipboard

Challenge: Korean pretrained language models struggle to generate short sentences with a given condition based on compositionality and commonsense reasoning.
Approach: They propose a Korean text-generation dataset for Korean generative commonsense reasoning and language model evaluation using a semi-automatic dataset construction approach.
Outcome: The proposed dataset is available at http://aihub.or.kr/opendata/korea-university.
RSGT: Relational Structure Guided Temporal Relation Extraction (2022.coling-1)

Copied to clipboard

Challenge: Temporal relation extraction (TRE) is crucial for natural language understanding.
Approach: They propose a Temporal Relational Structure Guided Temporal Relations Extraction task to extract relational structure features that can fit for both inter-sentence and intra-sentent relations.
Outcome: The proposed method improves on two well-known datasets, MATRES and TB-Dense, and can be used for clinical diagnosis and summarization.
Back to the Future: Bidirectional Information Decoupling Network for Multi-turn Dialogue Modeling (2022.emnlp-main)

Copied to clipboard

Challenge: Existing studies on dialogue modeling use pre-trained language models to encode dialogue history as successive tokens, which is insufficient in capturing the temporal characteristics of dialogues.
Approach: They propose a bidirectional information decoupling network as a universal dialogue encoder which explicitly incorporates both the past and future contexts.
Outcome: The proposed model incorporates past and future contexts and can be generalized to a wide range of dialogue-related tasks.
Probing Linguistic Systematicity (2020.acl-main)

Copied to clipboard

Challenge: Existing evidence that deep natural language understanding models do not learn systematically is lacking.
Approach: They examine whether deep natural language understanding models exhibit systematicity . they find that network architectures can generalize non-systematically .
Outcome: The proposed model generalizes non-systematically, but is unsatisfactory, the authors argue . they show that the current state-of-the-art models do not generalize systematically .
On Training Data Influence of GPT Models (2024.emnlp-main)

Copied to clipboard

Challenge: generative language models have redefined performance standards across tasks . current research on the influence of training data on autoregressivity remains underexplored .
Approach: They propose a parameterized simulation to assess the impact of training examples on the training dynamics of GPT models.
Outcome: The proposed approach compares existing methods with existing methods across training scenarios in generative language models, spanning tasks across 14 million to 2.8 billion parameters.
PARSE: LLM Driven Schema Optimization for Reliable Entity Extraction (2025.emnlp-industry)

Copied to clipboard

Challenge: Structured information extraction from unstructured text is critical for Software 3.0 systems . current approaches to extract structured information from unstructed text are static contracts .
Approach: They propose a system that automates JSON schemas for LLM consumption and optimizes them for LRM consumption.
Outcome: The proposed system improves extraction accuracy and reduces errors by 92% within the first retry and maintaining practical latency.
Visual Attention Model for Name Tagging in Multimodal Social Media (P18-1)

Copied to clipboard

Challenge: Name tagging is a key task for language understanding, but is often limited by the short textual components.
Approach: They propose a novel model architecture based on visual attention that outperforms other methods . they use multimodal datasets to analyze the name tagging task on social media .
Outcome: The proposed model outperforms existing methods and significantly outperformed existing methods.
It’s Not Just Size That Matters: Small Language Models Are Also Few-Shot Learners (2021.naacl-main)

Copied to clipboard

Challenge: Pretraining ever-larger language models on massive corpora requires enormous amounts of compute.
Approach: They propose to convert textual inputs into cloze questions that contain a task description . they also exploit unlabeled data to improve their performance .
Outcome: The proposed model outperforms GPT-3 with PET/iPET with cloze questions and unlabeled data.
Towards preserving word order importance through Forced Invalidation (2023.eacl-main)

Copied to clipboard

Challenge: Recent studies show pre-trained language models are insensitive to word order . performance on NLU tasks remains unchanged even after permuting the word .
Approach: They propose a simple approach called Forced Invalidation to force the model to identify permuted sequences as invalid samples.
Outcome: The proposed approach significantly improves the sensitivity of the models to word order on English NLU and QA tasks over BERT-based and attention-based models over word embeddings.
Program Synthesis for Complex QA on Charts via Probabilistic Grammar Based Filtered Iterative Back-Translation (2023.findings-eacl)

Copied to clipboard

Challenge: Current chart-based Question Answering approaches address structural, visual or simple data retrieval-type questions with fixed-vocabulary answers.
Approach: They employ a neural semantic parser to transform NL questions into SQL programs . they use a probabilistic context-free grammar to generate NL queries from a schema .
Outcome: The proposed approach achieves State-of-the-Art (SOTA) results on reasoning-based queries.
Meta-Reflection: A Feedback-Free Reflection Learning Framework (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to improve large language models' ability to understand and reason are limited by external feedback.
Approach: They propose a feedback-free reflection mechanism that requires only a single inference pass without external feedback.
Outcome: The proposed method is based on an industrial e-commerce benchmark and public datasets.
WYWEB: A NLP Evaluation Benchmark For Classical Chinese (2023.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for classical Chinese are inadequate to evaluate performance of different NLP models.
Approach: They propose an evaluation benchmark for classical Chinese NLP, which evaluates existing models.
Outcome: The proposed benchmark evaluates the performance of existing models in classical Chinese.
Efficient Large-Scale Neural Domain Classification with Personalized Attention (P18-1)

Copied to clipboard

Challenge: Using a scalable neural model, we show that personalization improves domain classification accuracy in a setting with thousands of overlapping domains.
Approach: They propose a scalable neural model architecture with a shared encoder that incorporates personalization information and domain-specific classifiers that solves the problem efficiently.
Outcome: The proposed architecture achieves two orders of magnitude faster than full model retraining.
Knowledge Prompting in Pre-trained Language Model for Natural Language Understanding (2022.emnlp-main)

Copied to clipboard

Challenge: Existing knowledge-enhanced pre-trained language models (PLMs) introduce redundant factual knowledge from knowledge bases and require complex modules.
Approach: They propose a knowledge prompting-based PLM framework that incorporates factual knowledge into PLMs.
Outcome: The proposed framework can be flexibly combined with existing mainstream PLMs.
Hyperparameter-free Continuous Learning for Domain Classification in Natural Language Understanding (2021.naacl-main)

Copied to clipboard

Challenge: Existing continual learning approaches suffer from low accuracy and performance fluctuation when the distributions of old and new data are significantly different.
Approach: They propose a hyperparameter-free continual learning model for text data that can stably produce high performance under various environments.
Outcome: The proposed model outperforms the best state-of-the-art method by 20% in average accuracy and each component contributes effectively to overall performance.
Unlocking Smarter Device Control: Foresighted Planning with a World Model-Driven Code Execution Approach (2025.findings-emnlp)

Copied to clipboard

Challenge: Current approaches to automating complex tasks focus on reactive policies and focus on visual observations.
Approach: They propose a framework that prioritizes natural language understanding and structured reasoning to enhance the agent’s global understanding of the environment by developing a task-oriented, refinable world model at the outset of the task.
Outcome: The proposed framework outperforms existing approaches in simulated environments and on real mobile devices.
Learning Numeracy: A Simple Yet Effective Number Embedding Approach Using Knowledge Graph (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing models for numeracy-intensive applications fail to learn numerability . existing models fail to handle numbers, resulting in performance problems .
Approach: They propose a number embedding approach that embeds numbers into dimensional space . they construct a knowledge graph consisting of number entities and magnitude relations .
Outcome: The proposed method is easy to implement and shows that it performs well on numeracy-related tasks.
LSOIE: A Large-Scale Dataset for Supervised Open Information Extraction (2021.eacl-main)

Copied to clipboard

Challenge: Open Information Extraction (OIE) systems extract factual propositions into n-ary tuples . current datasets are limited in size and diversity .
Approach: They propose to convert QA-SRL 2.0 dataset to large-scale OIE dataset LSOIE.
Outcome: The proposed dataset is 20 times larger than the next largest human-annotated OIE dataset.
Learning Global Controller in Latent Space for Parameter-Efficient Fine-Tuning (2024.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have shown remarkable performance, but their training costs are exorbitant.
Approach: They propose a parameter-efficient method for exploring optimal solutions within latent space by using latent units to extract input representations from LLMs.
Outcome: The proposed method improves performance on a range of natural language processing tasks.
MAVEN-ARG: Completing the Puzzle of All-in-One Event Understanding Dataset with Event Argument Annotation (2024.acl-long)

Copied to clipboard

Challenge: Existing datasets for event understanding have limited coverage due to complexity of tasks.
Approach: They propose a dataset that augments MAVEN datasets with event argument annotations . they propose 98,591 events and 290,613 arguments obtained with laborious human annotation .
Outcome: The proposed dataset is the first all-in-one dataset supporting event detection, event argument extraction, and event relation extraction.
QADiscourse - Discourse Relations as QA Pairs: Representation, Crowdsourcing and Baselines (2020.emnlp-main)

Copied to clipboard

Challenge: Discourse relations describe how two propositions relate to one another . annotating discourse relations requires expert annotators .
Approach: They propose a new representation of discourse relations as question-and-answer pairs that crowd-sources wide-coverage data annotated with discourse relations.
Outcome: The proposed representation of discourse relations as QA pairs allows crowd-sourcing wide-coverage datasets annotated with discourse relations.
Enhancing Reinforcement Learning with Label-Sensitive Reward for Natural Language Understanding (2024.acl-long)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have yielded remarkable performance, but objective mismatch issues hinder RLHF learning.
Approach: They propose a Reinforcement Learning framework enhanced with Label-sensitive reward to enhance LLMs' alignment and generation capabilities.
Outcome: The proposed framework improves performance on five diverse models across eight tasks.
Evaluating LLMs’ Mathematical Reasoning in Financial Document Question Answering (2024.findings-acl)

Copied to clipboard

Challenge: Large Language Models excel in natural language understanding, but their capability for complex mathematical reasoning with a hybrid of structured tables and unstructured text remain uncertain.
Approach: They propose a prompting technique tailored to semi-structured documents that matches or outperforms baselines performance while providing a nuanced understanding of LLMs' abilities.
Outcome: The proposed prompting technique outperforms baseline prompting techniques while providing a nuanced understanding of LLMs' abilities.
Improving Commonsense Question Answering by Graph-based Iterative Retrieval over Multiple Knowledge Sources (2020.coling-main)

Copied to clipboard

Challenge: Existing methods to facilitate natural language understanding rarely involve commonsense or background knowledge.
Approach: They propose a question-answering method that integrates multiple knowledge sources to boost performance.
Outcome: The proposed method outperforms other competing methods on the CommonsenseQA dataset and achieves the new state-of-the-art.
Rethinking Denoised Auto-Encoding in Language Pre-Training (2021.emnlp-main)

Copied to clipboard

Challenge: Pre-trained models such as BERT have achieved success in learning sequence representations, but they tend to learn representations that are covariant with the noise of pre-training.
Approach: They propose to train self-trained models to learn noise invariant sequence representations . they encourage consistency between original sequence and corrupted version via unsupervised instance-wise training signals.
Outcome: The proposed model improves on 11 natural language understanding and cross-modal tasks and achieves 0.6% gain on GLUE benchmarks and 0.8% increment on NLVR2 .
Curriculum: A Broad-Coverage Benchmark for Linguistic Phenomena in Natural Language Understanding (2022.naacl-main)

Copied to clipboard

Challenge: Existing evaluation methods do not provide insight into how well a language model captures distinct linguistic skills essential for language understanding and reasoning.
Approach: They propose a new format of NLI benchmark for evaluation of broad-coverage linguistic phenomena using a set of datasets and an evaluation procedure for diagnosing how well a language model captures reasoning skills.
Outcome: The proposed model can diagnose model behavior and verify model learning quality.
Ambiguity Meets Uncertainty: Investigating Uncertainty Estimation for Word Sense Disambiguation (2023.findings-acl)

Copied to clipboard

Challenge: Existing supervised methods treat word sense disambiguation as a classification task but ignore uncertainty estimation (UE) in the real-world setting, the data is always noisy and out of distribution.
Approach: They propose to use word sense disambiguation to determine an appropriate sense for a word given its context to determine the most appropriate sense.
Outcome: The proposed model reflects data uncertainty satisfactorily but underestimates model uncertainty.
Explainable Slot Type Attentions to Improve Joint Intent Detection and Slot Filling (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing methods analyze and compute features collectively for all slot types, and have no way to explain slot filling model decisions.
Approach: They propose a method that learns to generate additional slot type specific features to improve accuracy and provides explanations for slot filling decisions for the first time in a joint NLU model.
Outcome: The proposed model improves on two widely used datasets and provides an explanation for slot filling decisions for the first time.
Conversational Machine Comprehension: a Literature Review (2020.coling-main)

Copied to clipboard

Challenge: Conversational machine comprehension (CMC) is a research track in conversational AI.
Approach: They propose to synthesize a generic framework for CMC models and highlight differences in recent approaches.
Outcome: The proposed model will be used as a compendium for future research.
Natural Language Embedded Programs for Hybrid Language Symbolic Reasoning (2024.findings-naacl)

Copied to clipboard

Challenge: Existing methods for surfacing symbolic reasoning capabilities are limited to narrow tasks . arithmetic computations are unnatural to perform in pure language space, and hence present difficulties for LLMs.
Approach: They propose a natural language embedded program framework for solving symbolic reasoning tasks.
Outcome: The proposed framework improves on strong baselines across math and symbolic reasoning, text classification, question answering, and instruction following tasks.
Large Language Models are Temporal and Causal Reasoners for Video Question Answering (2023.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown remarkable performances on a wide range of natural language understanding and generation tasks.
Approach: They propose a framework that exploits linguistic shortcuts and mitigates 'linguistic bias' by flipping the source pair and target label to understand their complex relationships.
Outcome: The proposed framework outperforms both LLMs-based and non-LLMs- based models on five challenging VideoQA benchmarks.
ViGLUE: A Vietnamese General Language Understanding Benchmark and Analysis of Vietnamese Language Models (2024.findings-naacl)

Copied to clipboard

Challenge: Existing benchmarks for natural language understanding have been suggested, but there is a lack of such a benchmark in Vietnamese due to the difficulty in accessing datasets or the scarcity of task-specific datasets.
Approach: They propose to use a benchmark to evaluate Vietnamese language models in a variety of tasks and areas to explore the relationship between specific tasks and the number of shots.
Outcome: The proposed benchmark contains twelve tasks and encompasses over ten areas and subjects, enabling it to evaluate models comprehensively over a broad spectrum of aspects.
Syntax-BERT: Improving Pre-trained Transformers with Syntax Trees (2021.eacl-main)

Copied to clipboard

Challenge: Pre-trained language models like BERT achieve superior performances in various NLP tasks without explicit consideration of syntactic information.
Approach: They propose a plug-and-play framework that incorporates syntax trees into pre-trained Transformers.
Outcome: The proposed framework improves on pre-trained models on natural language understanding datasets and shows that it can be used to train pre-structured neural networks.
Inspecting the concept knowledge graph encoded by modern language models (2021.findings-acl)

Copied to clipboard

Challenge: Pre-trained language models are used to solve tasks such as summarization and information retrieval.
Approach: They propose to use word embeddings, text generators, context encoders to extract underlying knowledge graphs of nine influential language models.
Outcome: The proposed model is able to encode word embeddings, text generators, and context encoders, but suffers from several inaccuracies.
Plug-in and Fine-tuning: Bridging the Gap between Small Language Models and Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are renowned for their extensive linguistic knowledge and strong generalization capabilities, but their high computational demands make them unsuitable for resource-constrained environments.
Approach: They propose a framework that integrates a single frozen layer from an LLM into a SLM and fine-tunes the combined model for specific tasks.
Outcome: The proposed framework improves performance across a range of natural language processing tasks, including both natural language understanding and generation.
Semantically-Aligned Equation Generation for Solving and Reasoning Math Word Problems (N19-1)

Copied to clipboard

Challenge: Existing methods to solve math word problems require accurate natural language understanding to bridge texts and math expressions.
Approach: They propose a neural approach to automatically solve math word problems by operating symbols according to their semantic meanings in texts.
Outcome: The proposed model outperforms state-of-the-art models and the best non-retrieval-based models over 10% accuracy in a Math23K dataset.
No Context Needed: Contextual Quandary In Idiomatic Reasoning With Pre-Trained Language Models (2024.naacl-long)

Copied to clipboard

Challenge: idiomatic expressions (IEs) are a non-compositional aspect of a text that makes it difficult for a model to comprehend . general purpose PTLMs are negatively affected by the context, as performance increases with its removal.
Approach: They propose to use idiomatic expressions to infer additional meaning from IEs . they argue that only IE-aware models are suitable for idiom- matic reasoning tasks .
Outcome: The proposed models can reason in the presence of idiomatic expressions, the authors show . they show that general purpose PTLMs are negatively affected by the context .
CLOWER: A Pre-trained Language Model with Contrastive Learning over Word and Character Representations (2022.coling-1)

Copied to clipboard

Challenge: Pre-trained language models (PLMs) have achieved remarkable performance gains across numerous downstream tasks in natural language understanding.
Approach: They propose a Chinese pre-trained language model that implicitly encodes words into characters . they propose 'contrastive learning over word' and 'character' representations to improve learning .
Outcome: The proposed model can encode words into fine-grained representations without modification of production pipelines.
Out-of-Distribution Generalization in Natural Language Processing: Past, Present, and Future (2023.emnlp-main)

Copied to clipboard

Challenge: Existing literature on the generalization of machine learning models to out-of-distribution data is lacking.
Approach: They propose to present the first comprehensive review of recent progress, methods, and evaluations on the generalization challenge from an OOD perspective in natural language understanding.
Outcome: The proposed survey provides the first comprehensive review of recent progress, methods, and evaluations on the generalization challenge from an OOD perspective in natural language understanding.
MERIt: Meta-Path Guided Contrastive Learning for Logical Reasoning (2022.findings-acl)

Copied to clipboard

Challenge: Existing methods to infer logical relations with annotated training data suffer from over-fitting and poor generalization problems due to the dataset sparsity.
Approach: They propose a MEta-path guided contrastive learning method for logical ReasonIng of text that performs self-supervised pre-training on abundant unlabeled text data.
Outcome: The proposed method outperforms the baselines on two logical reasoning benchmarks with significant improvements.
AutoLoRA: Automatically Tuning Matrix Ranks in Low-Rank Adaptation Based on Meta Learning (2024.naacl-long)

Copied to clipboard

Challenge: Large-scale pretraining followed by task-specific finetuning has achieved great success in various NLP tasks.
Approach: They propose a meta learning based framework for automatically identifying the optimal rank of each LoRA layer.
Outcome: The proposed framework is based on a meta learning based framework that can identify the optimal rank of each LoRA layer.
ConFiguRe: Exploring Discourse-level Chinese Figures of Speech (2022.coling-1)

Copied to clipboard

Challenge: Figures of speech often deviate from their literal meanings to express deeper semantic implications.
Approach: They propose a concept of figurative unit, which is the carrier of a figure, and build a Chinese corpus for Contextualized Figure Recognition.
Outcome: The proposed model is based on 12 types of figures commonly used in Chinese . it shows that the proposed tasks are challenging for existing models .
Statistically Profiling Biases in Natural Language Reasoning Datasets and Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to evaluate NLP models' weaknesses are limited by “hypothesis-only” tests and CheckLists.
Approach: They propose a lightweight general statistical profiling framework that automatically identifies potential biases in multiple-choice NLU datasets without requiring additional test cases.
Outcome: The proposed framework assesses the extent to which models exploit these biases through black-box testing, confirming prior findings and revealing new insights.
ExPUNations: Augmenting Puns with Keywords and Explanations (2022.emnlp-main)

Copied to clipboard

Challenge: Puns add the challenge of fusing commonsense and world knowledge with the ability to interpret lexical-semantic ambiguity.
Approach: They propose to augment existing datasets with detailed crowdsourced annotations of puns, keywords and fine-grained funniness ratings to challenge current models' ability to understand and generate humor.
Outcome: The proposed tasks include explanation generation to aid with pun classification and keyword-conditioned pun generation to challenge state-of-the-art models' ability to understand and generate humor.
A Context-based Approach for Dialogue Act Recognition using Simple Recurrent Neural Networks (L18-1)

Copied to clipboard

Challenge: Existing models of dialogue act classification work on the utterance-level and only very few consider context.
Approach: They propose to use a character-level language model to classify dialogue acts without context . they find that the preceding utterances are a context of the current utterant .
Outcome: The proposed method improves on the Switchboard Dialogue Act corpus . it includes context and leads to 3% higher accuracy .
Towards an Automatic Assessment of Crowdsourced Data for NLU (L18-1)

Copied to clipboard

Challenge: Recent development of spoken dialog systems aims at allowing a natural input style.
Approach: They investigate how crowdsourced data can be assessed with respect to its naturalness and usefulness by using a word based language model to identify valid data.
Outcome: The proposed methods show that valid data can be identified with the help of a word based language model.
NLEBench+NorGLM: A Comprehensive Empirical Analysis and Benchmark Dataset for Generative Language Models in Norwegian (2024.emnlp-main)

Copied to clipboard

Challenge: Norwegian is under-represented within the most impressive breakthroughs in NLP tasks.
Approach: they investigate the impact of existing Norwegian language models on Norwegian generation tasks . they pre-trained 4 Norwegian Open Language Models from parameter scales and architectures .
Outcome: The proposed benchmark evaluates the performance of language models on Norwegian generation tasks.
JGLUE: Japanese General Language Understanding Evaluation (2022.lrec-1)

Copied to clipboard

Challenge: There is no benchmark for Japanese to evaluate and analyze NLU ability from different perspectives.
Approach: They build a Japanese NLU benchmark from scratch without translation to measure general NLU ability in Japanese.
Outcome: a Japanese NLU benchmark is built from scratch without translation to measure general NLU ability in Japanese.
Causal Distillation for Language Models (2022.naacl-main)

Copied to clipboard

Challenge: Distillation efforts have led to language models that are more compact and efficient without serious drops in performance.
Approach: They propose to augment distillation with a third objective that encourages the student model to imitate the causal dynamics of the teacher through a distillation interchange intervention training objective (DIITO).
Outcome: The proposed method lowers perplexity on the WikiText-103M corpus and improves on the GLUE benchmark, SQuAD, and CoNLL-2003.
A recipe for annotating grounded clarifications (2021.naacl-main)

Copied to clipboard

Challenge: In order to interpret communicative intents of an utterance, it needs to be grounded in world modalities.
Approach: They propose a recipe for obtaining grounding annotations for dialogue clarification mechanisms that make explicit the process of interpreting communicative intents of an utterance.
Outcome: The proposed method is based on the definitions of perceptual and collaborative grounding and on the classification of clarification phenomena.
Tiny Budgets, Big Gains: Parameter Placement Strategy in Parameter Super-Efficient Fine-Tuning (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods such as LoRA and VeRA use memory-efficient methods to fine-tune large language models.
Approach: They propose a method that uses only 1–5% of the standard LoRA parameters and achieves state-of-the-art performance across a wide range of tasks.
Outcome: The proposed method achieves state-of-the-art performance across a wide range of tasks using only 1–5% of the standard LoRA parameters.
Verb Metaphor Detection via Contextual Relation Learning (2021.acl-long)

Copied to clipboard

Challenge: Recent work on verb metaphor detection focuses on analyzing restricted forms of linguistic context.
Approach: They propose a model which explicitly models the relation between a verb and its various contexts.
Outcome: The proposed model gets competitive results compared with state-of-the-art approaches on the VUA, MOH-X and TroFi datasets.
GeoHard: Towards Measuring Class-wise Hardness through Modelling Class Semantics (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in measuring hardness-wise properties of data guide language models in sample selection within low-resource scenarios.
Approach: They propose to use class-wise hardness to measure class-specific properties of data in the semantic embedding space by modeling class geometry in the . semantic embeddining space.
Outcome: The proposed method surpasses instance-level metrics by over 59 percent on Pearson‘s correlation on measuring class-wise hardness.
SOCCER: An Information-Sparse Discourse State Tracking Collection in the Sports Commentary Domain (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods for state tracking are limited and state changes are less densely distributed over utterances.
Approach: They propose to turn to simplified, fully observable systems that show some of these properties.
Outcome: The proposed system shows that state changes occur infrequently while messages are "chatter" it allows for rich descriptions of state while avoiding the complexities of other settings.
Learning Explainable Linguistic Expressions with Neural Inductive Logic Programming for Sentence Classification (2020.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to explain models are difficult to interpret and have undesirable biases.
Approach: They propose a neural network architecture for learning transparent sentences . they use linguistic expressions built on top of predicates extracted using shallow natural language understanding .
Outcome: The proposed model outperforms statistical relational learning and other neuro-symbolic methods and performs better than black-box recurrent neural networks.
Adaptation with Self-Evaluation to Improve Selective Prediction in LLMs (2023.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have shown impressive capabilities in many tasks, including natural language understanding and generation.
Approach: They propose a framework for adaptation with self-evaluation to improve selective prediction performance of large language models.
Outcome: The proposed framework outperforms state-of-the-art selective prediction methods on QA datasets and improves the AUACC from 91.23% to 92.63% and AUROC from 74.61% to 80.25%.
Bi-Drop: Enhancing Fine-tuning Generalization via Synchronous sub-net Estimation and Optimization (2023.findings-emnlp)

Copied to clipboard

Challenge: Pretrained language models can be fine-tuned on limited training data, which can overfit and thus diminish performance.
Approach: They propose a fine-tuning strategy that selectively updates model parameters using gradients from various sub-nets dynamically generated by dropout.
Outcome: The proposed method outperforms existing methods on the GLUE benchmark and exhibits excellent generalization ability and robustness for domain transfer, data imbalance, and low-resource scenarios.
Effective Large Language Model Adaptation for Improved Grounding and Citation Generation (2024.naacl-long)

Copied to clipboard

Challenge: Large language models generate "hallucinated" answers that are not factual . despite their widespread adoption, they can generate plausiblesounding but nonfactual information.
Approach: They propose a framework that tunes large language models to self-ground claims and provide citations to retrieved documents.
Outcome: The proposed framework generates superior grounded responses with more accurate citations compared to prompting-based approaches and post-hoc citing-based methods.
FaVIQ: FAct Verification from Information-seeking Questions (2022.acl-long)

Copied to clipboard

Challenge: Existing fact verification datasets with crowdsourced claims introduce subtle biases that are difficult to control for.
Approach: They construct a large-scale fact verification dataset with ambiguous questions . they use a corpus of 188k claims to construct false and true claims .
Outcome: The proposed dataset outperforms models trained on the dataset FEVER or in-domain data by up to 17% absolute.
MiniELM: A Lightweight and Adaptive Query Rewriting Framework for E-Commerce Search Optimization (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for rewriting query terms struggle with natural language understanding . generative methods face high inference latency and cost in offline settings .
Approach: They propose a hybrid pipeline for rewriting query queries using offline knowledge distillation and online reinforcement learning.
Outcome: The proposed pipeline improves query relevance, diversity, adaptability and cost-effective evaluation without manual annotations on Amazon ESCI dataset.
On the Calibration of Pre-trained Language Models using Mixup Guided by Area Under the Margin and Saliency (2022.acl-long)

Copied to clipboard

Challenge: Existing studies have shown that mixing up can improve model calibration on image classification tasks, but little is known about using it on natural language understanding (NLU) tasks.
Approach: They propose a mixup strategy for pre-trained language models that improves model calibration further by using the AUM statistic and saliency map.
Outcome: The proposed mixup improves model calibration on natural language understanding tasks while maintaining competitive accuracy.
Tiny-NewsRec: Effective and Efficient PLM-based News Recommendation (2022.emnlp-main)

Copied to clipboard

Challenge: Existing work fine tunes the PLM with the news recommendation task, which can cause a domain shift problem.
Approach: They propose a self-supervised method to adapt general PLM to news domain with a contrastive matching task between news titles and news bodies.
Outcome: The proposed method can improve both the effectiveness and efficiency of the large PLM-based news recommendation model while maintaining its performance.
Cross-Lingual NLU: Mitigating Language-Specific Impact in Embeddings Leveraging Adversarial Learning (2024.lrec-main)

Copied to clipboard

Challenge: Low-resource languages and computational expenses pose significant challenges in the domain of large language models.
Approach: They propose a novel approach that uses adversarial techniques to mitigate the impact of language-specific information in contextual embeddings generated by large multilingual language models.
Outcome: The proposed approach excels in zero-shot scenarios for Latin languages like Spanish, but fails to perform for languages distant from English, such as Thai and Persian.
Sequential Cross-Document Coreference Resolution (2021.emnlp-main)

Copied to clipboard

Challenge: Existing models for cross-document coreference resolution have been used for within-document entity coreference but have been relatively limited.
Approach: They propose a model that extends the efficient sequential prediction paradigm for coreference resolution to cross-document settings and achieves competitive results for both entity and event coreference.
Outcome: The proposed model achieves competitive results for entity and event coreference while minimizing error propagation in complex reasoning tasks.
What Will it Take to Fix Benchmarking in Natural Language Understanding? (2021.naacl-main)

Copied to clipboard

Challenge: Evaluation for many natural language understanding (NLU) tasks is broken due to unreliable and biased systems scoring so high on standard benchmarks.
Approach: They argue that current benchmarks fail at four criteria for evaluation . they argue that adversarial data collection does not address the causes of failures .
Outcome: The proposed frameworks fail at four criteria, and adversarial data collection does not address the causes of these failures, the authors argue . restoring a healthy evaluation ecosystem will require significant progress in the design of benchmark datasets, reliability with which they are annotated, their size, and the ways they handle social bias.
From text to talk: Harnessing conversational corpora for humane and diversity-aware language technology (2022.acl-long)

Copied to clipboard

Challenge: Informal social interaction is the primordial home of human language.
Approach: They show that linguistically diverse conversational corpora can provide empirical foundations for flexible, localizable language technologies of the future.
Outcome: The results suggest that even relatively small corpora can support robust generalizations about key aspects of interactional infrastructure.
TreeMix: Compositional Constituency-based Data Augmentation for Natural Language Understanding (2022.naacl-main)

Copied to clipboard

Challenge: Existing data augmentation methods miss the important characteristic of compositionality, meaning of a complex expression is built from its sub-parts.
Approach: They propose a compositional data augmentation approach for natural language understanding called TreeMix that leverages constituency parsing tree to decompose sentences into constituent sub-structures and the Mixup data enhancing technique to recombine them to generate new sentences.
Outcome: The proposed approach outperforms current state-of-the-art methods on text classification and SCAN.
Continuation KD: Improved Knowledge Distillation through the Lens of Continuation Optimization (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for knowledge distillation (KD) do not mitigate the noise in the teacher’s output: modeling the noisy behaviour of the teacher can distract the student from learning more useful features.
Approach: They propose a method that optimizes the highly non-convex KD objective by starting with the smoothed version of this objective and making it more complex as the training proceeds.
Outcome: The proposed method achieves state-of-the-art performance on NLU and computer vision tasks.
RoMe: A Robust Metric for Evaluating Natural Language Generation (2022.acl-long)

Copied to clipboard

Challenge: Empirical results suggest that RoMe has a stronger correlation to human judgment over state-of-the-art metrics in evaluating system-generated sentences across several NLG tasks.
Approach: They propose an automatic evaluation metric incorporating several core aspects of natural language understanding (language competence, syntactic and semantic variation).
Outcome: The proposed evaluation metric is trained on language features such as semantic similarity combined with tree edit distance and grammatical acceptability, using a self-supervised neural network.
Pay Attention when you Pay the Bills. A Multilingual Corpus with Dependency-based and Semantic Annotation of Collocations. (P19-1)

Copied to clipboard

Challenge: resulting corpus can be useful for different NLP tasks such as natural language understanding or natural language generation.
Approach: They propose to annotate 155k tokens and 1,526 collocations in context in a multilingual corpus in English, Portuguese, and Spanish.
Outcome: The new corpus can be used to evaluate different approaches for collocation identification, which can be useful for different NLP tasks such as natural language understanding or natural language generation.
Self-Distillation for Model Stacking Unlocks Cross-Lingual NLU in 200+ Languages (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) excel on English NLU tasks, yet struggle to extend their NLU capabilities to underrepresented languages.
Approach: They integrate machine translation models (MT) directly into LLM backbones via sample-efficient self-distillation.
Outcome: The proposed model outperforms translation-test models on 127 low-resource languages.
InforMask: Unsupervised Informative Masking for Language Model Pretraining (2022.emnlp-main)

Copied to clipboard

Challenge: Masked language modeling is used for pretraining large language models for knowledge-intensive tasks.
Approach: They propose an unsupervised masking strategy that exploits Pointwise Mutual Information to select the most informative tokens to mask.
Outcome: The proposed strategy outperforms random masking and previously proposed masking strategies on the factual recall benchmark LAMA and the question answering benchmark SQuAD v1 and v2.
Key ingredients for effective zero-shot cross-lingual knowledge transfer in generative tasks (2024.naacl-long)

Copied to clipboard

Challenge: Existing studies have focused on zero-shot cross-lingual transfer . mBERT, mBART and mT5 provide high-quality representations for texts in various languages .
Approach: They propose to use mBART and NLLB-200 to finetune a multilingual pretrained language model on input-output pairs in one language and use it to make task predictions for inputs in other languages.
Outcome: The proposed approach significantly reduces generation in the wrong language with full finetuning and can be competitive in some cases.
Prompt-Tuning Can Be Much Better Than Fine-Tuning on Cross-lingual Understanding With Multilingual Language Models (2022.findings-emnlp)

Copied to clipboard

Challenge: Pre-trained multilingual language models show significant performance gains for zero-shot cross-lingual model transfer on a wide range of natural language understanding (NLU) tasks.
Approach: They do cross-lingual evaluation using prompt tuning and compare it with fine-tuning . prompt tuning achieves much better cross-linguistic transfer than fine- tuning .
Outcome: The results show that prompt tuning achieves better cross-lingual transfer than fine-tuning across datasets, with only 0.1% to 0.3% tuned parameters.
CLUE: A Chinese Language Understanding Evaluation Benchmark (2020.coling-main)

Copied to clipboard

Challenge: Existing language evaluation benchmarks for English are limited to English . lack of such benchmarks makes it difficult to replicate success in other languages .
Approach: They introduce a large-scale Chinese language understanding evaluation benchmark . the benchmark uses a set of current state-of-the-art pre-trained Chinese models .
Outcome: The first large-scale Chinese Language Understanding Evaluation (CLUE) benchmark is released . the benchmark evaluates models across a wide range of tasks on original Chinese text . existing language evaluation benchmarks are mostly limited to English .
SenseBERT: Driving Some Sense into BERT (2020.acl-main)

Copied to clipboard

Challenge: Existing approaches for self-supervision operate at word form level, which serves as a surrogate for the underlying semantic content.
Approach: They propose a method to employ weak-supervision directly at the word sense level, without the use of human annotation.
Outcome: The proposed model achieves significantly improved lexical understanding without human annotation on the ‘Word in Context’ task.
Self-training Improves Pre-training for Natural Language Understanding (2021.naacl-main)

Copied to clipboard

Challenge: Unsupervised pretraining has led to improvements in natural language understanding . a data augmentation method can be used to generate labels for unlabeled examples .
Approach: They propose a semi-supervised method which uses unlabeled data to retrieve sentences from a database of billions of unlabed sentences crawled from the web.
Outcome: The proposed method improves on standard text classification benchmarks by 2.6% and knowledge distillation by few shots.
WikiCREM: A Large Unsupervised Corpus for Coreference Resolution (D19-1)

Copied to clipboard

Challenge: Large-scale training sets for pronoun resolution are scarce, since manually labelling data is costly.
Approach: They propose a language-model-based approach to solve pronoun disambiguation problems using a WikiCREM dataset.
Outcome: The proposed model outperforms state-of-the-art approaches on 6 out of 7 datasets.
Multi-Task Deep Neural Networks for Natural Language Understanding (P19-1)

Copied to clipboard

Challenge: Existing approaches to learning vector-space representations of text are multitask learning and language model pre-training.
Approach: They propose a multi-task deep neural network (MT-DNN) that leverages cross-task data and incorporates a pre-trained bidirectional transformer language model.
Outcome: The proposed model achieves state-of-the-art on ten NLU tasks and pushes the GLUE benchmark to 82.7% (2.2% absolute improvement)
DisSent: Learning Sentence Representations from Explicit Discourse Relations (P19-1)

Copied to clipboard

Challenge: Existing models train on vast amounts of text or require costly, manually curated datasets.
Approach: They propose to leverage the discourse relations between sentences to curate a high quality sentence relation task by leveraging explicit discourse relations.
Outcome: The proposed model can be used to learn the meaning of two sentences in a bidirectional LSTM sentence encoder.
Coreference Reasoning in Machine Reading Comprehension (2021.acl-long)

Copied to clipboard

Challenge: Existing datasets for machine reading comprehension do not reflect the natural distribution and, consequently, the challenges of coreference reasoning.
Approach: They propose to use existing coreference resolution datasets to train machine reading comprehension models to better reflect the natural distribution and, consequently, the challenges of coreference reasoning.
Outcome: The proposed method improves the performance of state-of-the-art models on a set of coreference-related datasets.
Hierarchical Transformer for Task Oriented Dialog Systems (2021.naacl-main)

Copied to clipboard

Challenge: Existing models for dialog generation are challenging to train using the standard Seq2Seq models.
Approach: They propose a framework for Hierarchical Transformer Encoders that can be morphed into any hierarchical transformer by using specially designed attention masks and positional encodings.
Outcome: The proposed framework can be morphed into any hierarchical encoder, including HRED and HIBERT like models, by using specially designed attention masks and positional encodings.
CLUTRR: A Diagnostic Benchmark for Inductive Reasoning from Text (D19-1)

Copied to clipboard

Challenge: Existing datasets for reading comprehension tasks have been used to test the generalization of natural language understanding systems.
Approach: They propose a diagnostic benchmark suite to clarify key issues related to the robustness and systematicity of NLU systems.
Outcome: The proposed benchmark suite clarifies key issues related to the robustness and systematicity of NLU systems.
HeGeL: A Novel Dataset for Geo-Location from Hebrew Text (2023.findings-acl)

Copied to clipboard

Challenge: Existing datasets in English for textual geolocation are limited because of the location of the place is implicit.
Approach: They propose to use a Hebrew place description corpus to analyze lingual geospatial reasoning.
Outcome: The Hebrew Geo-Location corpus collects literal Hebrew place descriptions and analyzes lingual geospatial reasoning.
Identifying Motion Entities in Natural Language and A Case Study for Named Entity Recognition (2020.coling-main)

Copied to clipboard

Challenge: Identifying motion entities in text is not only challenging but beneficial for a better natural language understanding.
Approach: They propose a Motion Entity Tagging model to identify entities in motion in a text using the Literal-Motion-in-Text dataset for training and evaluating the model.
Outcome: The proposed method improves the Named-Entity Recognition task by splitting clauses and phrases from complex and long motion sentences.
Universal Self-Adaptive Prompting (2023.emnlp-main)

Copied to clipboard

Challenge: a hallmark of modern large language models is their impressive general zero-shot and few-shot abilities . however, zero- shot performances are weaker due to the lack of guidance and the difficulty of applying existing automatic prompt design methods in general tasks.
Approach: They propose an automatic prompt design approach specifically tailored for zero-shot learning that categorizes a possible NLP task into one of three possible task types and then uses a selector to select the most suitable queries and zero- shot model-generated responses as pseudo-demonstrations.
Outcome: The proposed approach is able to generalize ICL to zero-shot learning tasks while also allowing for a more efficient and efficient prompt design.
(Re)construing Meaning in NLP (2020.acl-main)

Copied to clipboard

Challenge: a new paper explores the role of linguistic choices in interpreting information in natural language . linguistic choice is a way of expressing information, but it is not the meaning of an utterance, authors argue .
Approach: They propose to define construal as a way of conceptualizing or construing information . they propose to use this concept to develop theoretical and practical work in NLP .
Outcome: The proposed study explores how construal can inform theoretical and practical work in NLP.
Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data (2020.acl-main)

Copied to clipboard

Challenge: a priori, large neural language models are described as understanding or capturing meaning on tasks that are ostensibly meaningsensitive.
Approach: They argue that a system trained only on form has no way to learn meaning . they argue that this is due to a misunderstanding of the relationship between form and meaning - which is a misconception in NLP .
Outcome: The proposed model can't learn meaning because it only uses form as training data, the authors argue . they argue that a clear understanding of the distinction between form and meaning will guide the field towards better science around natural language understanding.
How Can We Accelerate Progress Towards Human-like Linguistic Generalization? (2020.acl-main)

Copied to clipboard

Challenge: a new evaluation paradigm, Pretraining-Agnostic Identically Distributed evaluation, is needed . authors argue that it rewards models that can be trained on massive amounts of data, several orders of magnitude more than a human can expect to be exposed to.
Approach: a position paper describes and critiques the Pretraining-Agnostic Identically Distributed evaluation paradigm . paradigm favors simple, low-bias architectures that can be scaled to process vast amounts of data . authors advocate for supplementing or replacing PAID with paradigms that reward architectures .
Outcome: a new evaluation paradigm favors simple, low-bias architectures that can be scaled to process vast amounts of data. a san francisco-based study finds that the paradigm rewards architectures which generalize as quickly and robustly as humans.
Lost in Inference: Rediscovering the Role of Natural Language Inference for Large Language Models (2025.naacl-long)

Copied to clipboard

Challenge: In the recent past, a popular way of evaluating natural language understanding was to consider a model’s ability to perform natural language inference (NLI) tasks.
Approach: They focus on five different NLI benchmarks across six models of different scales and examine how their accuracies develop during training.
Outcome: The softmax distributions of models align with human label distributions in cases where statements are ambiguous or vague.
Would you Rather? A New Benchmark for Learning Machine Alignment with Cultural Values and Social Preferences (2020.acl-main)

Copied to clipboard

Challenge: Existing studies on optimal decision-making are limited and only consider individuals in isolation.
Approach: They propose a task and corpus for learning alignments between machine and human preferences based on a gamified voting game .
Outcome: The proposed task and corpus show that current state-of-the-art NLP models still leave much room for improvement.
A Surprisingly Robust Trick for the Winograd Schema Challenge (P19-1)

Copied to clipboard

Challenge: The Winograd Schema Challenge (WSC) dataset WSC273 and its inference counterpart WNLI are popular benchmarks for natural language understanding and commonsense reasoning.
Approach: They propose to fine-tune language models on the Winograd Schema Challenge dataset WSC273 and its inference counterpart WNLI to achieve accuracies of 72.5% and 74.7%, respectively.
Outcome: The proposed language models achieve 72.5% and 74.7% accuracy on the WSC273 and WNLI datasets, respectively.
What Makes Reading Comprehension Questions Difficult? (2022.acl-long)

Copied to clipboard

Challenge: a recent study shows that natural language understanding benchmarks are not able to measure future progress . a crowdsourcing approach is needed to collect diverse examples without sacrificing diversity or coverage.
Approach: They crowdsource multiple-choice reading comprehension questions for passages from seven sources . they find passage source, length, and readability measures do not significantly affect question difficulty .
Outcome: The results show that passage source, length, and readability measures do not significantly affect question difficulty.
Cross-Lingual Training for Automatic Question Generation (P19-1)

Copied to clipboard

Challenge: Automatic question generation is a challenging problem in natural language understanding . manual curating a dataset of comparable size for a new language is tedious and expensive.
Approach: They propose to reuse available large QG dataset in a secondary language to learn a QG model for a primary language.
Outcome: The proposed model outperforms baseline models in Hindi and Chinese.
XGLUE: A New Benchmark Dataset for Cross-lingual Pre-training, Understanding and Generation (2020.emnlp-main)

Copied to clipboard

Challenge: XGLUE provides a benchmark dataset to train large-scale cross-lingual pre-trained models . XCLUE provides 11 diversified tasks that cover both understanding and generation scenarios .
Approach: They introduce a new benchmark dataset to train large-scale cross-lingual pre-trained models using multilingual and bilingual corpora.
Outcome: The proposed dataset is labeled in English and includes only natural language understanding tasks.
Towards Debiasing Sentence Representations (2020.acl-main)

Copied to clipboard

Challenge: Recent work has shown word-level embeddings reflect and propagate social biases present in training corpora.
Approach: They propose a method to debias word embeddings to reduce biases at sentence level . they hope their work will inspire future research on characterizing and removing biase .
Outcome: The proposed method reduces biases and preserves performance on downstream tasks such as sentiment analysis and natural language understanding.
Parameter-Efficient Tuning Makes a Good Classification Head (2022.emnlp-main)

Copied to clipboard

Challenge: In recent years, pretrained models revolutionized the paradigm of natural language understanding . but the final-layer output of the backbone, i.e. the input of the classification head, will change greatly during finetuning .
Approach: They propose to append a randomly initialized classification head after the pretrained backbone and finetune the whole model.
Outcome: The proposed classification head can be replaced with the randomly initialized heads for a stable performance gain.
An Attribution Relations Corpus for Political News (L18-1)

Copied to clipboard

Challenge: Existing resources for recognizing attributions in context are limited in size and completeness.
Approach: They propose to use the largest and most complete attribution relations corpus to date . they propose to create sophisticated end-to-end solutions for attribution extraction .
Outcome: The political news attribution relations corpus 2016 is the largest and most complete attribution relations corpuse to date.
ZeroSCROLLS: A Zero-Shot Benchmark for Long Text Understanding (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmarks for long text understanding focus on short sequences, such as BigBench and HELM.
Approach: They propose a zero-shot benchmark for natural language understanding over long texts . they adapt six tasks from the SCROLLS benchmark and add four new datasets .
Outcome: The proposed benchmark outperforms ChatGPT and GPT-4 in a number of open tasks.
Curriculum Learning for Natural Language Understanding (2020.acl-main)

Copied to clipboard

Challenge: Pre-trained language models can be fine tuned to perform NLU tasks in a straightforward manner.
Approach: They propose a pretrain-finetune paradigm for natural language understanding (NLU) they propose 'a cross-trainset' approach that allows users to distinguish easy from difficult examples .
Outcome: The proposed approach achieves significant performance improvements on a wide range of NLU tasks.
Dual Supervised Learning for Natural Language Understanding and Generation (P19-1)

Copied to clipboard

Challenge: Natural language understanding and natural language generation are important research topics in the NLP and dialogue fields.
Approach: They propose a dual-supervised learning framework for natural language understanding and generation on top of dual supervised learning.
Outcome: The proposed framework boosts the performance of both tasks simultaneously in the benchmark experiments.
Question Directed Graph Attention Network for Numerical Reasoning over Text (2020.emnlp-main)

Copied to clipboard

Challenge: Numerical reasoning requires both natural language understanding and arithmetic computation.
Approach: They propose a graph representation for the context of the passage and question needed for numerical reasoning.
Outcome: The proposed model achieves remarkable results in benchmark datasets such as DROP.
MindMap: Knowledge Graph Prompting Sparks Graph of Thoughts in Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Large language models suffer from limitations such as difficulty in incorporating new knowledge, generating hallucinations, and explaining their reasoning process.
Approach: They propose a pipeline that leverages knowledge graphs to enhance LLMs’ inference and transparency by eliciting the mind map of LLM's, which reveals their reasoning pathways based on the ontology of knowledge.
Outcome: The proposed pipeline enables LLMs to comprehend KG inputs and infer with a combination of implicit and external knowledge.
The Unstoppable Rise of Computational Linguistics in Deep Learning (2020.acl-main)

Copied to clipboard

Challenge: a quarter century ago, linguists assumed that language knowledge needed to be innate . but vector-space representations and machine learning algorithms are much more powerful than was thought .
Approach: They trace the history of neural networks applied to natural language understanding tasks . they argue that Transformer is not a sequence model but an induced-structure model .
Outcome: The proposed model is not a sequence model but an induced-structure model, the authors argue . they argue that the nature of language has had a profound impact on progress in machine learning .
A Top-down Neural Architecture towards Text-level Parsing of Discourse Rhetorical Structure (2020.acl-main)

Copied to clipboard

Challenge: Text-level discourse parsing of discourse rhetorical structure (DRS) is a fundamental research topic in natural language processing.
Approach: They propose a top-down neural architecture for text-level discourse parsing . they cast the parser as a recursive split point ranking task .
Outcome: The proposed top-down approach is more suitable for text-level discourse parsing.
Predicting Humorousness and Metaphor Novelty with Gaussian Process Preference Learning (P19-1)

Copied to clipboard

Challenge: Inability to quantify key aspects of creative language is a frequent obstacle to natural language understanding.
Approach: They propose a Bayesian approach for predicting humorousness and metaphor novelty using Gaussian process preference learning (GPPL) .
Outcome: The proposed approach achieves a Spearman’s of 0.56 against gold using word embeddings and linguistic features.
Pruning Pre-trained Language Models with Principled Importance and Self-regularization (2023.findings-acl)

Copied to clipboard

Challenge: Pre-trained language models often contain a vast amount of parameters, posing nontrivial requirements for storage and computation.
Approach: They propose a pruning method where model prediction is regularized by the latest checkpoint with increasing sparsity throughout pruning.
Outcome: The proposed approach is effective at sparsity levels, and can be applied to natural language understanding, question answering, and data-to-text generation tasks.
Words Aren’t Enough, Their Order Matters: On the Robustness of Grounding Visual Referring Expressions (2020.acl-main)

Copied to clipboard

Challenge: Visual referring expression recognition is a task that requires natural language understanding in the context of an image.
Approach: They propose to use contrastive learning and multi-task learning to increase the robustness of ViLBERT, the current state-of-the-art model for this task.
Outcome: The proposed methods are 12% to 23% lower in performance than the established progress for this task.
StraGo: Harnessing Strategic Guidance for Prompt Optimization (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for prompt optimization often lead to prompt drifting, wherein newly generated prompts canadversely impact previously successful cases while addressing failures.
Approach: They propose a method to mitigate prompt drifting by integrating in-context learning to formulate specific, actionable strategies for prompt optimization.
Outcome: The proposed approach mitigates prompt drifting by leveraging insights from both successful and failed cases to identify critical factors for achieving optimization objectives.
MetaPro 2.0: Computational Metaphor Processing on the Effectiveness of Anomalous Language Modeling (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods for metaphor interpretation are slow due to lack of annotated datasets and effective pre-trained language models.
Approach: They propose a large annotated dataset and a PLM for the metaphor interpretation task.
Outcome: The proposed method improves on metaphor identification and interpretation with comparable baselines on the new dataset.
tagE: Enabling an Embodied Agent to Understand Human Instructions (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing systems for natural language understanding (NLU) are limited due to the inherent ambiguity and incompleteness inherent in natural language.
Approach: They propose a system to extract tasks from natural language instructions and map them to robots' established collection of skills.
Outcome: The proposed system outperforms baseline models in the training and evaluation of a dataset featuring complex instructions.
Two Birds One Stone: Dynamic Ensemble for OOD Intent Classification (2023.acl-long)

Copied to clipboard

Challenge: Out-of-domain (OOD) intent classification is an active field of natural language understanding . previous studies have suggested that PTMs would be "overthinking" the semantic features of the sample in the open-world scenario .
Approach: They propose a method that allows the model to decide whether to make a decision on OOD classification early during inference.
Outcome: The proposed method can improve inference speed and achieve significant performance improvements.
BERT Knows Punta Cana is not just beautiful, it’s gorgeous: Ranking Scalar Adjectives with Contextualised Representations (2020.emnlp-main)

Copied to clipboard

Challenge: Adjectives describe positive properties of nouns but with different intensity.
Approach: They propose a BERT-based approach to intensity detection for scalar adjectives by generating vectors directly from contextualised representations.
Outcome: The proposed model outperforms static embeddings and previous models with dedicated resources on an Indirect Question Answering task.
TURNA: A Turkish Encoder-Decoder Language Model for Enhanced Understanding and Generation (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in natural language processing have favored well-resourced English-centric models, resulting in a significant gap with low-resource languages.
Approach: They propose a language model for the low-resource language Turkish that is capable of both natural language understanding and generation tasks.
Outcome: The proposed model outperforms multilingual models in understanding and generation tasks and competes with monolingual models for understanding tasks.
RikiNet: Reading Wikipedia Pages for Natural Question Answering (2020.acl-main)

Copied to clipboard

Challenge: Using Wikipedia pages to answer open-domain questions remains challenging in natural language understanding.
Approach: They propose a model which reads Wikipedia pages for natural question answering . it uses a dynamic paragraph dual-attention reader and a cascaded answer predictor .
Outcome: The proposed model outperforms the human model on the Natural Questions dataset . it achieves 74.3 F1 and 57.9 F1 on long-answer and short-answer tasks .
MedCare: Advancing Medical LLMs through Decoupling Clinical Alignment and Knowledge Aggregation (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have made significant progress in natural language understanding and generation, proving valuable especially in the medical field.
Approach: They propose a medical LLM through decoupling Clinical Alignment and Knowledge Aggregation which uses a and a to encode diverse knowledge in the first stage and filter out detrimental information.
Outcome: The proposed model achieves promising performance on over 20 medical tasks and specific medical alignment tasks.
A Survey on LLM-powered Agents for Recommender Systems (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models have demonstrated remarkable capabilities in natural language understanding, reasoning, and generation.
Approach: They present a comprehensive synthesis of large language models and their applications . they dissect a four-module agent architecture and review representative designs .
Outcome: The proposed models address fundamental challenges in traditional recommender systems . they include limited comprehension of complex user intents, insufficient interaction capabilities .
Cross-Lingual Semantic Role Labeling with High-Quality Translated Training Corpus (2020.acl-main)

Copied to clipboard

Challenge: Existing approaches to semantic role labeling (SRL) are focusing on the English language.
Approach: They propose a method for semantic role labeling that uses corpus translation to build training datasets from SRL annotations.
Outcome: The proposed method is highly effective and can improve the target-language performance significantly.
Pipeline Analysis for Developing Instruct LLMs in Low-Resource Languages: A Case Study on Basque (2025.naacl-long)

Copied to clipboard

Challenge: Large language models are typically optimized for resource-rich languages like English . however, the proprietary nature of these models makes them impractical for many researchers and developers.
Approach: They propose to develop large language models that can follow instructions in Basque . they focus on three key stages: pre-training, instruction tuning, and alignment with human preferences .
Outcome: The proposed models improve natural language understanding (NLU) of the foundational model by 12 points . the results show that the models can follow instructions in Basque with human preferences .
tBERT: Topic Models and BERT Joining Forces for Semantic Similarity Detection (2020.acl-main)

Copied to clipboard

Challenge: Recent pretrained contextual representations such as ELMo and BERT have led to impressive performance gains across a variety of NLP tasks, including semantic similarity detection.
Approach: They propose a topic-informed BERT-based architecture for pairwise semantic similarity detection that adds topic information to pretrained contextual representations such as BERT.
Outcome: The proposed model outperforms existing models on a variety of English language datasets and is highly performant.
Extracting Event Temporal Relations via Hyperbolic Geometry (2021.emnlp-main)

Copied to clipboard

Challenge: Recent neural approaches to event temporal relation extraction map events to embeddings in the Euclidean space and train a classifier to detect temporal relations between event pairs.
Approach: They propose to embed events into hyperbolic spaces to model hierarchical structures . they propose to use hyperbolical embeddings to directly infer event relations .
Outcome: The proposed architecture is based on two approaches to encode events and their temporal relations in hyperbolic spaces.
A Large Interlinked Knowledge Graph of the Italian Cultural Heritage (2022.lrec-1)

Copied to clipboard

Challenge: Existing efforts to create knowledge bases are limited to relatively small resources, such as entities from libraries, archeological sites and museums.
Approach: They propose to create a large knowledge graph linking Italian cultural heritage entities with concepts defined on well-known knowledge bases.
Outcome: The proposed graph shows that the Italian cultural heritage entities are interlinked with concepts defined on well-known knowledge bases.
Understanding Programs by Exploiting (Fuzzing) Test Cases (2023.findings-acl)

Copied to clipboard

Challenge: Semantic understanding of programs has attracted great attention in the community . large language models (LLMs) are capable of learning contextual information from data at scale .
Approach: They propose to incorporate a relationship between inputs and possible outputs into learning for achieving a deeper semantic understanding of programs.
Outcome: The proposed method outperforms current state-of-the-art on two programming tasks and outperformed current state of the art by large margins.
PALM: Pre-training an Autoencoding&Autoregressive Language Model for Context-conditioned Generation (2020.emnlp-main)

Copied to clipboard

Challenge: Existing techniques for natural language understanding and generation use autoencoding and/or autoregressive objectives to train models.
Approach: They propose a self-supervised pre-training scheme that pre-trains an autoencoding and autoregressive language model on a large unlabeled corpus for generating new text conditioned on context.
Outcome: The proposed scheme achieves state-of-the-art results on a variety of language generation benchmarks covering generative question answering, abstractive summarization and conversational response generation.
Creation of a Balanced State-of-the-Art Multilayer Corpus for NLU (L18-1)

Copied to clipboard

Challenge: Using full stack of language resources, we are creating a balanced text corpus for Latvian.
Approach: They propose to create a syntactically and semantically annotated multilayered corpus for Latvian . they use widely acknowledged and cross-lingual representations for the corpus .
Outcome: The proposed corpus adopts widely recognized and cross-lingual representations for natural language understanding and generation in Latvian.
From Spatial Relations to Spatial Configurations (2020.lrec-1)

Copied to clipboard

Challenge: Existing spatial representations are not sufficient for describing complex spatial configurations.
Approach: They propose to integrate existing spatial representation languages with an annotation schema to extend the capabilities of existing ones.
Outcome: The proposed language can represent a large set of spatial concepts crucial for reasoning . it integrates with the Abstract Meaning Representation (AMR) annotation schema and annotates text from diverse datasets .
Debias NLU Datasets via Training-free Perturbations (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to debiase NLU models capture biased features that are independent of the task but spuriously correlated to labels.
Approach: They propose a framework that conducts training-free perturbations on samples containing biased features to Debias NLU Datasets.
Outcome: The proposed framework shows competitive performance with previous state-of-the-art debiasing strategies.
Disfluency Generation for More Robust Dialogue Systems (2023.findings-acl)

Copied to clipboard

Challenge: Disfluencies in user utterances can trigger a chain of errors impacting all the modules of a dialogue system.
Approach: They propose to augment existing dialogue datasets with disfluent utterances by paraphrasing them into disfluente ones.
Outcome: The proposed method improves dialogue state tracking and response generation by combining disfluent utterances with disfluency utteraces.
AdapterShare: Task Correlation Modeling with Adapter Differentiation (2022.emnlp-main)

Copied to clipboard

Challenge: AdapterShare is an adapter differentiation method to explicitly model the task correlation among multiple tasks.
Approach: They propose an adapter differentiation method to explicitly model the task correlation among multiple tasks.
Outcome: The proposed method achieves 1.90 points improvement on five dialogue understanding tasks and 2.33 points gain on NLU tasks.
InternalInspector I2: Robust Confidence Estimation in LLMs through Internal States (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) often struggle with generating reliable outputs, often producing high-confidence inaccuracies known as hallucinations.
Approach: They propose a framework that leverages contrastive learning on internal states including attention states, feed-forward states, and activation states of all layers to enhance confidence estimation in LLMs.
Outcome: The framework outperforms existing methods in the hallucination detection benchmark HaluEval and the previous methods at the same time.
Are Natural Language Inference Models IMPPRESsive? Learning IMPlicature and PRESupposition (2020.acl-main)

Copied to clipboard

Challenge: Natural language inference (NLI) is an increasingly important task for natural language understanding . however, the ability of NLI models to make pragmatic inferences remains understudied .
Approach: They use semi-automatically generated sentence pairs to evaluate whether NLI models make pragmatic inferences.
Outcome: The proposed model trains on multiNLI and shows that it learns to draw pragmatic inferences.
Mind the Trade-off: Debiasing NLU Models without Degrading the In-distribution Performance (2020.acl-main)

Copied to clipboard

Challenge: Recent studies show that pre-trained language models rely heavily on idiosyncratic biases of datasets.
Approach: They propose a method which discourages models from exploiting biases while enabling them to receive enough incentive to learn from all the training examples.
Outcome: The proposed method improves on out-of-distribution datasets while maintaining original in-district accuracy.
Instance Regularization for Discriminative Language Model Pre-training (2022.emnlp-main)

Copied to clipboard

Challenge: Existing studies have optimized independent strategies of ennoising or denosing . Existing methods treat training instances equally throughout the training process .
Approach: They propose to use ennoising and denoising to train discriminative pre-trained language models . they propose to model the complexity of restoring the original sentences from corrupted ones .
Outcome: Experimental results show that the proposed method improves pre-training efficiency, effectiveness, and robustness.
Explaining Interactions Between Text Spans (2023.emnlp-main)

Copied to clipboard

Challenge: Existing highlight-based explanations focus on identifying individual important features or interactions only between adjacent tokens or tuples of tokens.
Approach: They propose a multi-annotator dataset of human span interaction explanations for NLU and FC.
Outcome: The proposed method compares human reasoning processes to those of a fine-tuned large language model.
When More Data Hurts: A Troubling Quirk in Developing Broad-Coverage Natural Language Understanding Systems (2022.emnlp-main)

Copied to clipboard

Challenge: In natural language understanding systems, users’ evolving needs necessitate the addition of new features over time, indexed by new symbols added to the meaning representation space.
Approach: They propose to use a small set of new symbols to build broad-coverage NLU systems.
Outcome: The proposed model is based on two prototypical NLU tasks: intent recognition and semantic parsing.
Exploring Reversal Mathematical Reasoning Ability for Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have been a success in the wide range of natural language understanding and reasoning tasks.
Approach: They propose a training method to improve general and reversal reasoning abilities by using a reversed dataset.
Outcome: The proposed method improves general and reversal reasoning abilities and alleviates the reverse curse.
IT5: Text-to-text Pretraining for Italian Language Understanding and Generation (2024.lrec-main)

Copied to clipboard

Challenge: Xue et al., 2022) use the text-to-text paradigm to train multilingual models.
Approach: They introduce the first family of encoder-decoder transformer models pretrain specifically on Italian and introduce the ItaGen benchmark to evaluate the models' performance.
Outcome: The proposed model outperforms models with multilingual baselines and the original model on English data.
TextLap: Customizing Language Models for Text-to-Layout Planning (2024.findings-emnlp)

Copied to clipboard

Challenge: Creating 2D graphical layouts from text alone is challenging in traditional settings.
Approach: They propose to customize LLMs to allow users to generate professional looking layouts by simply inputting text instructions.
Outcome: The proposed method outperforms existing benchmarks for document generation and graphical design benchmarks.
Training Multi-Modal LLMs through Dialogue Planning for HRI (2025.findings-acl)

Copied to clipboard

Challenge: Existing approaches to enhance Multi-Modal Large Language Models (MLLMs) with explicit dialogue planning improves response accuracy and quality, and allows models trained in one language to transfer effectively to another.
Approach: They propose an approach that enhances Multi-Modal Large Language Models with a novel explicit dialogue planning phase that allows agents to refine their understanding of ambiguous commands.
Outcome: The proposed approach reduces hallucinations and improves task feasibility by fine-tuning and assessing Multi-Modal models in human-robot interaction scenarios.
Decoder Tuning: Efficient Language Understanding as Decoding (2023.acl-long)

Copied to clipboard

Challenge: Existing approaches to adapt pre-trained models with parameters frozen are based on input-side adaptation, which requires thousands of API queries.
Approach: They propose to train a model-as-a-service (MaaS) setting to provide only the inference APIs for users . they argue that input-side adaptation could be arduous due to the lack of gradient signals .
Outcome: The proposed model outperforms state-of-the-art algorithms with a 200x speed-up.
The KITMUS Test: Evaluating Knowledge Integration from Multiple Sources (2023.acl-long)

Copied to clipboard

Challenge: Existing models that make inferences using information from multiple sources are largely understudied .
Approach: They propose a test suite of coreference resolution subtasks that require reasoning over multiple facts and introduce subtask where knowledge is present only at inference time using fictional knowledge.
Outcome: The proposed subtasks differ in terms of which knowledge sources contain the relevant facts and where knowledge is present only at inference time using fictional knowledge.
Looking Right is Sometimes Right: Investigating the Capabilities of Decoder-only LLMs for Sequence Labeling (2024.findings-acl)

Copied to clipboard

Challenge: Pre-trained language models excel in natural language understanding (NLU) tasks.
Approach: They propose to apply layer-dependent removal of the causal mask (CM) during LLM fine-tuning to improve SL performance.
Outcome: The proposed approach outperforms state-of-the-art SL models on IE tasks, while achieving state- of-the art results is unclear.
RanLoRA: Residual-aware Nonlinear Low-Rank Adaptation (2026.findings-acl)

Copied to clipboard

Challenge: Low-Rank Adaptation (LoRA) relying on linear low-rank projections restricts adaptation to linear subspaces, limiting flexibility on complex downstream tasks.
Approach: They propose a nonlinear low-rank Adaptation approach that leverages pretrained weights to decompose them into principal components that are kept frozen and residual components that can be used for task-specific adaptation.
Outcome: The proposed approach outperforms vanilla LoRA and representative variants on commonsense reasoning, image classification, and mathematical reasoning tasks.
NusaCrowd: Open Source Initiative for Indonesian NLP Resources (2023.findings-acl)

Copied to clipboard

Challenge: Existing NLP research in Indonesian languages has been held back by factors such as language diversity, orthographic variation, resource limitation and other societal challenges.
Approach: They present a collaborative initiative to collect and unify existing resources for Indonesian languages and open access to previously non-public resources.
Outcome: The results show that the datasets are highly reliable and can be used to generate the first zero-shot benchmarks for natural language understanding and generation in Indonesian and the local languages of Indonesia.
HyperMixer: An MLP-based Low Cost Alternative to Transformers (2023.acl-long)

Copied to clipboard

Challenge: Existing MLP-based architectures that combine multiple features are expensive and require a lot of training data.
Approach: They propose a simple MLP-based model which allows token mixing by dynamically applying hypernetworks to each feature independently.
Outcome: The proposed model performs better than Transformers and lowers costs in terms of processing time, training data, and hyperparameter tuning.
TLM: Token-Level Masking for Transformers (2023.emnlp-main)

Copied to clipboard

Challenge: Structured dropout approaches have been investigated to regularize the multi-head attention mechanism in Transformers.
Approach: They propose a new regularization scheme based on token-level rather than structure-level to reduce overfitting by manipulating the connections between tokens in the multi-head attention via masking.
Outcome: The proposed regularization scheme outperforms attention dropout and DropHead on 18 datasets and can establish a new record on the data-to-text benchmark Rotowire (18.93 BLEU).
IEKG: A Commonsense Knowledge Graph for Idiomatic Expressions (2023.emnlp-main)

Copied to clipboard

Challenge: Prior work on IE comprehension has focused on detecting idiomaticity, but this fails to account for IEs' non-compositionality.
Approach: They construct a commonsense knowledge graph for figurative interpretations of IEs that can be used to convert PTLMs into knowledge models that encode and infer commonsensical knowledge related to IE use.
Outcome: The proposed model can generalize to IEs unseen during training.
Layer-wise Regularized Dropout for Neural Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods to regularize dropout are consistency training and dropout is a problem in many pre-trained neural language models.
Approach: They propose a layer-wise regularized dropout technique which regularizes dropout at the output layer using consistency training.
Outcome: The proposed model can be regarded as a "self-distillation" framework, in which each sub-model generated by dropout is the other's "teacher" model and "student" model.
Natural Logic at the Core: Dynamic Rewards for Entailment Tree Generation (2025.findings-acl)

Copied to clipboard

Challenge: Existing approaches to generating entailment trees lack logical consistency . static reward structures or intricate dependencies within multi-step reasoning are often ignored .
Approach: They propose a method that integrates natural logic principles into reinforcement learning to guide entailment tree generation.
Outcome: Experiments on EntailmentBank show that the proposed method improves interpretability and generalization.
ACCEPT: Adaptive Codebook for Composite and Efficient Prompt Tuning (2024.findings-emnlp)

Copied to clipboard

Challenge: Prompt Tuning has been a popular fine-tuning method for large-scale pretrained language models.
Approach: They propose a method that allows all soft prompts to share a set of learnable codebook vectors in each subspace, with each prompt differentiated by a number of adaptive weights.
Outcome: The proposed method achieves superior performance on 17 diverse natural language tasks including natural language understanding (NLU) and question answering (QA) tasks by tuning only 0.3% of parameters of the PLMs.
HeQ: a Large and Diverse Hebrew Reading Comprehension Benchmark (2023.findings-emnlp)

Copied to clipboard

Challenge: Current benchmarks for Hebrew Natural Language Processing (NLP) focus mainly on morpho-syntactic tasks, neglecting the semantic dimension of language understanding.
Approach: They propose to use Hebrew machine reading comprehension (MRC) as extractive Question Answering to address this problem.
Outcome: The proposed benchmark features 30,147 question-answer pairs derived from both Hebrew Wikipedia articles and Israeli tech news.
Unveiling the Essence of Poetry: Introducing a Comprehensive Dataset and Benchmark for Poem Summarization (2023.emnlp-main)

Copied to clipboard

Challenge: Summarization of poetry is a challenging task as it can be easily lost if only the literal meaning is considered.
Approach: They propose to use poetry as a model to summarize poetry and provide a dataset to evaluate their creative language interpretation capacity.
Outcome: The proposed dataset consisting of 3011 samples and its corresponding summarized interpretation in the English language provides an opportunity to evaluate the creative language interpretation capacity of the proposed models.
It Ain’t Over: A Multi-aspect Diverse Math Word Problem Dataset (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies lack diversity in problem types, lexical usage patterns, languages, and intermediate solution forms for the math word problem.
Approach: They propose a new MWP dataset with a wide range of diversity in problem types, lexical usage patterns, languages, and intermediate solutions.
Outcome: The proposed dataset provides an opportunity to evaluate the capability of large language models.
Make Prompt-based Black-Box Tuning Colorful: Boosting Model Generalization from Three Orthogonal Perspectives (2024.lrec-main)

Copied to clipboard

Challenge: Large language models (LLMs) have shown increasing power on NLP tasks. however, tuning these models for downstream tasks usually requires exorbitant costs.
Approach: They propose a black-box tuning technique that optimizes task-specific prompts without accessing gradients and hidden representations.
Outcome: The proposed method improves performance under few-shot learning scenarios.
Mitigating Shortcuts in Language Models with Soft Label Encoding (2024.lrec-main)

Copied to clipboard

Challenge: Recent studies have shown that large language models rely on spurious correlations in the data for natural language understanding (NLU) tasks.
Approach: They propose a framework for debiasing shortcuts and a dummy class to encode shortcuts into a model and use it to generate soft labels.
Outcome: The proposed framework significantly improves out-of-distribution generalization while maintaining satisfactory in-district accuracy.
Model-Agnostic Cross-Lingual Training for Discourse Representation Structure Parsing (2024.lrec-main)

Copied to clipboard

Challenge: Discourse Representation Structure (DRS) parsers are constrained when trained exclusively on monolingual data.
Approach: They propose a cross-lingual training strategy that leverages cross-linguistic training data to train models in multiple languages.
Outcome: The proposed method improves clause and graph parsing in English, German, Italian and Dutch.
MaCP: Minimal yet Mighty Adaptation via Hierarchical Cosine Projection (2025.acl-long)

Copied to clipboard

Challenge: MaCP is a new adaptation method for large foundation models that requires minimal parameters and memory for fine-tuning.
Approach: They propose a method that exploits the superior energy compaction and decorrelation properties of cosine projection to improve model efficiency and accuracy.
Outcome: The proposed method improves model efficiency and accuracy across a wide range of single-modality tasks including natural language understanding, natural language generation, text summarization, and multi-modalities such as image classification and video understanding.
Contrastive Pre-training for Personalized Expert Finding (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to expert finding are effective for a community question answering platform.
Approach: They propose a CQA-domain Contrastive Pre-training framework for Expert Finding which could learn more comprehensive question representations.
Outcome: The proposed framework could learn more comprehensive question representations on six real-world datasets.
VerifyMatch: A Semi-Supervised Learning Paradigm for Natural Language Inference with Confidence-Aware MixUp (2024.emnlp-main)

Copied to clipboard

Challenge: Natural language inference (NLI) is a key task for evaluating a model's ability to perform natural language understanding and reasoning.
Approach: They propose to construct pseudo-generated samples using class-specific fine-tuned large language models (LLMs) . they retain all pseudo-labeled samples, but use MixUp to ensure unlabele .
Outcome: The proposed approach achieves competitive accuracy compared to strong baselines for NLI datasets in low-resource settings.
Benchmarking Large Language Models for Cryptanalysis and Side-Channel Vulnerabilities (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have transformed natural language understanding and generation, leading to extensive benchmarking across diverse tasks.
Approach: They evaluate the cryptanalytic potential of stateoftheart LLMs on ciphertexts produced by a range of cryptographic algorithms.
Outcome: The proposed model can decrypt plaintexts produced by a range of cryptographic algorithms using zeroshot and fewshot settings along with chainofthought prompting.
Exploring Explanations Improves the Robustness of In-Context Learning (2025.acl-long)

Copied to clipboard

Challenge: In-context learning (ICL) has been shown to be effective across a variety of tasks, but it has been reported to be restricted in its ability to generalize beyond the given demonstrations.
Approach: They propose a framework that extends ICL by exploring explanations for all possible labels.
Outcome: The proposed framework improves prediction reliability by exploring explanations for all possible labels.
Prompting Large Language Models for Counterfactual Generation: An Empirical Study (2024.lrec-main)

Copied to clipboard

Challenge: Large language models (LLMs) have made remarkable progress in a wide range of natural language understanding and generation tasks, but their ability to generate counterfactuals has not been examined systematically.
Approach: They propose a framework to evaluate LLMs' ability to generate counterfactuals based on key factors including intrinsic properties and prompt design.
Outcome: The proposed framework examines the strengths and weaknesses of large language models (LLMs) and identifies factors that influence their ability to generate counterfactuals.
RRAtention: Dynamic Block Sparse Attention via Per-Head Round-Robin Shifts for Long-Context Inference (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to dynamic sparse attention require preprocessing, lack global evaluation, violate query independence, or incur high computational overhead.
Approach: They propose a dynamic sparse attention method that achieves all desirable properties through a head **r**ound-**r**obin (RR) sampling strategy.
Outcome: Experiments on natural language understanding and multimodal video comprehension show that the proposed method achieves 2.4 speedup at 128K context length outperforming existing methods.
Revisiting Data Reconstruction Attacks on Real-world Dataset for Federated Natural Language Understanding (2024.lrec-main)

Copied to clipboard

Challenge: Existing DRA methods fail to accurately recover the original text of real-world privacy data.
Approach: They propose to use a real-world privacy dataset to examine the performance of federated learning (FL) methods.
Outcome: The proposed method improves on a real-world privacy dataset and shows that the tokens within a recovery sentence are disordered and intertwined with tokens from other sentences in the same training batch.
RoCoIns: Enhancing Robustness of Large Language Models through Code-Style Instructions (2024.lrec-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown remarkable capabilities in following human instructions and solving NLU tasks.
Approach: They propose to use code style instructions to replace typically natural language instructions to provide more precise instructions and strengthen the robustness of LLMs.
Outcome: The proposed method outperforms natural language models on eight robustness datasets and achieves an improvement of 5.68% in test set accuracy and a reduction of 5.66 points in Attack Success Rate (ASR).
Large Language Models Meet Knowledge Graphs for Question Answering: Synthesis and Opportunities (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have shown remarkable performance on question-answering tasks due to their superior capabilities in natural language understanding and generation.
Approach: They propose a structured taxonomy that categorizes the methodology of synthesizing LLMs and knowledge graphs for QA according to the categories of QA and the KG’s role when integrating with LLM.
Outcome: The proposed taxonomy categorizes the methods according to the categories of QA and the KG’s role when integrating with LLMs.
Arctic-Text2SQL-R1: Simple Rewards, Strong Reasoning in Text-to-SQL (2026.findings-acl)

Copied to clipboard

Challenge: Translating natural language questions into SQL is a core challenge in natural language understanding and human-computer interaction.
Approach: They propose a reinforcement learning framework and model family to generate accurate, executable SQL using a lightweight reward signal based solely on execution correctness.
Outcome: The proposed framework outperforms previous versions of 70B-class systems and achieves state-of-the-art execution accuracy across six diverse Text2SQL benchmarks.
TLoRA: Task-aware Low Rank Adaptation of Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing low-rank Adaptation (LoRA) methods address only one factor, often at the cost of increased training complexity or reduced practical efficiency.
Approach: They propose a low-rank Adaptation framework that optimizes initialization and resource allocation at the outset of training.
Outcome: The proposed framework performs excellently across various tasks while reducing the number of trainable parameters.
Injecting Domain-Specific Knowledge into Large Language Models: A Comprehensive Survey (2025.findings-emnlp)

Copied to clipboard

Challenge: specialized LLMs are often limited in domain-specific applications that require specialized knowledge.
Approach: They provide a comprehensive overview of four key methods to enhance large language models by integrating domain-specific knowledge.
Outcome: The proposed methods are categorized into four key approaches: dynamic knowledge injection, static knowledge embedding, modular adapters, and prompt optimization.
Improving the Language Understanding Capabilities of Large Language Models Using Reinforcement Learning (2025.findings-emnlp)

Copied to clipboard

Challenge: Instruction-fine-tuned large language models (LLMs) under 14B parameters underperform on NLU tasks . we explore a framework to improve the NLU capabilities of LLMs .
Approach: They propose to use Proximal Policy Optimization to improve NLU capabilities . they frame NLU as a reinforcement learning environment and optimize for reward signals .
Outcome: The proposed framework outperforms supervised fine-tuning on GLUE and superGLUE tasks.
Preference Estimation via Opponent Modeling in Multi-Agent Negotiation (2026.findings-acl)

Copied to clipboard

Challenge: Existing numerical-only approaches fail to capture qualitative information embedded in natural language interactions, resulting in unstable and incomplete preference estimation.
Approach: They propose a preference estimation method that integrates natural language information into a Bayesian opponent modeling framework.
Outcome: The proposed method improves agreement rate and preference estimation accuracy by integrating probabilistic reasoning with natural language understanding.
Towards Equitable Natural Language Understanding Systems for Dialectal Cohorts: Debiasing Training Data (2024.lrec-main)

Copied to clipboard

Challenge: Prior research has shown that biases exist in these models against certain languages or dialects.
Approach: They propose to use a dialect identification model to obtain targeted training data augmentation for under-represented dialects to debias NLU model for dialectal cohorts in NLU systems.
Outcome: The proposed framework can provide insights on dialect disparity in real-world NLU systems and targeted data argumentation can help narrow the model’s performance gap between standard language speakers and dialect speakers.
Transformer-based Swedish Semantic Role Labeling through Transfer Learning (2024.lrec-main)

Copied to clipboard

Challenge: Semantic Role Labeling (SRL) is a task in natural language understanding where the goal is to extract semantic roles for a given sentence.
Approach: They propose to build a Transformer-based SRL system for Swedish by exploring multilingual and cross-lingual transfer learning methods and leveraging the Swedish FrameNet resource.
Outcome: The proposed model outperforms two different cross-lingual transfer models and shows that the multilingual learning outperformed the other models.
What Can Diachronic Contexts and Topics Tell Us about the Present-Day Compositionality of English Noun Compounds? (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods to determine the semantic relatedness between compounds and constituents have applied a synchronic perspective, but this study examines what diachronic changes in contexts and semantic topics reveal about the compounds’ present-day compositionality.
Approach: They propose to use two diachronic vector spaces to model compositional patterns between compounds with low and high present-day compositionality.
Outcome: The proposed model performs on par with co-occurrence space and captures similar information.
Astra: Activation-Space Tail-Eigenvector Low-Rank Adaptation of Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for fine-tuning pre-trained models are limited due to suboptimal activation subspaces.
Approach: They propose a method that leverages tail eigenvectors of model output activations to construct low-rank adapters.
Outcome: The proposed method outperforms existing methods across 16 benchmarks and surpasses full fine-tuning in certain scenarios.
Unlocking Human-Like Visible Logic: How Logic Diagrams Boost Logic Reasoning in Large Language Models? (2026.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have demonstrated their remarkable capabilities in natural language understanding and generation, but they struggle with formal logical reasoning.
Approach: They propose to incorporate visual logic diagrams into LLMs’ reasoning workflows to enhance their performance on formal logic tasks.
Outcome: The proposed model improves on syllogistic and conditional reasoning with programmatically generated Venn, Euler, and Linear diagrams.
CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding (2026.findings-acl)

Copied to clipboard

Challenge: Multimodal large language models generate medical hallucinations due to over-sensitivity to clinical sections.
Approach: They propose a framework that integrates structured clinical signals from task-specific radiology expert models.
Outcome: The proposed framework improves overall performance on radiology report generation (RRG) on the MIMIC-CXR dataset, it yields up to 17% improvement in RadGraph-F1.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations